ASRU 2025 · ARXIV 2506.13053 · 中文全文译稿

ZipVoice: Fast and High-Quality
Zero-Shot Text-to-Speech
with Flow Matching

Han Zhu · Wei Kang · Zengwei Yao · Liyong Guo · Fangjun Kuang · Zhaoqing Li · Weiji Zhuang · Long Lin · Daniel PoveyXiaomi Corp. · Beijing, China联系邮箱:{zhuhan3, dpovey}@xiaomi.com
01 / CONDITIONPhone Zipformer192 维文本状态,经平均上采样变成帧级条件
02 / GENERATORZipformer Flow123M 参数,预测条件流匹配向量场
03 / SAMPLER4 NFE Distill蒸馏 CFG 轨迹,推理不再做双分支前向
PDF
8 页
Figures
1
Tables
6
Equations
11
References
55

现有大规模 zero-shot text-to-speech(TTS)模型虽然能生成高质量语音,却因参数规模庞大而推理缓慢。为解决这一问题,本文提出 ZipVoice:一种基于 Flow Matching、模型紧凑且推理快速的高质量 zero-shot TTS。其关键设计包括:1)采用 Zipformer-based vector field estimator,在受限参数规模下维持足够的建模能力;2)用 Average Upsampling 建立初始语音—文本对齐,并通过 Zipformer-based text encoder 提升语音可懂度;3)提出 flow distillation,减少采样步数,同时消除 classifier-free guidance 带来的额外推理开销。作者在 10 万小时多语言数据上的实验表明,ZipVoice 的语音质量可与 SOTA 模型匹配,同时模型规模仅为 DiT-based Flow Matching baseline 的约三分之一,推理最高快 30 倍。代码、模型 checkpoint 与 demo 样例均已公开。

关键词: Text-to-speech、zero-shot、Flow Matching。

1 引言

Text-to-speech(TTS)的目标,是合成能够准确表达给定文本的自然语音。Zero-shot TTS 还要求生成语音模仿参考音频中的说话人特征,从而灵活适配任意声音。近年来,生成建模方法 [1]–[7] 与大规模数据集 [8]–[10] 推动了 zero-shot TTS 的显著进展。理想的 TTS 系统不仅应生成高质量语音,还应保持简单和高效。然而,现有在大规模数据上训练的 zero-shot TTS 模型通常依赖大量参数来维持足够的建模能力;多数 SOTA 系统还需要重复采样,例如随时间执行 autoregressive(AR)采样 [1],或采用 diffusion model 中的 non-autoregressive(NAR)采样 [11]。庞大的模型与大量重复采样步骤共同拖慢推理,提高部署成本,并限制 zero-shot TTS 的实际应用。

近来,作为 diffusion model 变体的 Flow Matching [12] 因采样步数少于早期 diffusion-based 方法,被广泛用于 TTS [2]–[4], [13], [14]。E2-TTS [3] 是其中具有代表性的 zero-shot TTS 系统:它用 filler token 补齐文本 token,使其长度与语音一致,再让基于 Transformer [15] 的 vector field estimator 隐式学习文本—语音对齐。由于不再需要 token-level duration prediction,这一做法显著简化了系统结构与训练流程。但 E2-TTS 仍需较大模型和数十次采样才能取得满意效果,因此推理依旧缓慢。

为解决上述问题,本文提出 ZipVoice:一个模型紧凑、推理快速的高质量 zero-shot TTS。ZipVoice 以结构简单且表现优良的 Flow Matching 为基础,并通过三项设计兼顾效率与质量。第一,使用 Zipformer [16] 作为 vector field estimator 的 backbone,在有限参数预算内保留充足建模能力;作者证明,Zipformer 虽最初为 automatic speech recognition(ASR)设计,也适合 TTS。第二,E2-TTS 的 filler-token padding 会造成次优的文本—语音对齐并损害可懂度 [4]。ZipVoice 一方面采用 Average Upsampling,假设同一句话中每个 token 的时长相同,在不引入 token-level duration predictor 的情况下提供稳定初始对齐;另一方面用 Zipformer-based text encoder 提取更好的 text-token representation。第三,Flow Matching TTS 通常仍需多步采样,而常用的 classifier-free guidance(CFG)[17] 又会在每一步增加一次 unconditional forward。作者因此提出 flow distillation,将预训练 Flow Matching 模型蒸馏为少步模型,并消除 CFG 的双倍推理开销。

实验结果表明:与 DiT-based Flow Matching baseline [4] 相比,ZipVoice 的模型规模约小 3 倍、推理最高快 30 倍,同时在可懂度、说话人相似度和自然度方面仍可与现有 SOTA zero-shot TTS 系统相当。

2 ZipVoice

本节介绍所提出的 zero-shot TTS 模型 ZipVoice。首先简要回顾 Flow Matching。

Figure 1:ZipVoice 的训练流程(左)与推理流程(右)。

随后给出 ZipVoice 概览,并详细说明两个以 Zipformer 为 backbone 的子模型;接着介绍 Average Upsampling;之后说明使用 flow distillation 训练的加速版本 ZipVoice-Distill;最后介绍模型的推理策略。

2.1 预备知识:Flow Matching

ZipVoice 建立在 Flow Matching 框架 [12] 上,具体采用 Conditional Flow Matching(CFM)。下面先简要介绍 CFM。

CFM 学习把简单初始分布 p0p_0(例如标准 Gaussian distribution)变换为逼近真实数据分布 qq 的复杂分布 p1p_1。模型由 time-dependent vector field vt(xt;θ),t[0,1]v_t(x_t; \theta), t\in[0,1] 参数化,并据此构造把 p0p_0 推向 p1p_1 的 flow ϕt\phi_t。在 optimal-transport 路径 ϕt(x)=(1t)x0+tx1\phi_t(x)=(1-t)x_0+tx_1 下,CFM 损失为:

LCFM=Et,q(x1),p0(x0)vt(xt;θ)(x1x0)2(1)L_{\text{CFM}}=E_{t,q(x_1),p_0(x_0)}\lVert v_t(x_t;\theta)-(x_1-x_0)\rVert^2\tag{1}

其中,xt=(1t)x0+tx1x_t=(1-t)x_0+tx_1x0x_0 为 Gaussian noise,x1x_1 为数据样本。

用 CFM 损失训练后,可通过求解 ordinary differential equation(ODE)生成样本。以 Euler solver 为例,从初始样本 x0p0x_0\sim p_0(如 Gaussian noise)出发,迭代地把样本推向目标分布 p1p_1。对于离散时间序列 0=t0<<tk<<tK=10=t_0<\dots<t_k<\dots<t_K=1,第 kk 步更新为:

xtk+1=xtk+(tk+1tk)vtk(xtk;θ)(2)x_{t_{k+1}}=x_{t_k}+(t_{k+1}-t_k)\cdot v_{t_k}(x_{t_k};\theta)\tag{2}

这一过程将 x0x_0 逐步变换为 x1p1x_1\sim p_1。Function evaluation 次数(NFE)KK 决定推理速度与样本质量之间的取舍。

Classifier-free guidance(CFG)[17] 常用于提高 Flow Matching 模型的生成质量。训练时以一定概率丢弃条件,推理时则对 conditional 与 unconditional prediction 做线性组合:

v~t(xt,c,ω;θ)=(1+ω)vt(xt,c;θ)ωvt(xt,;θ)(3)\begin{aligned}\tilde{v}_t(x_t,c,\omega;\theta)&=(1+\omega)v_t(x_t,c;\theta)\\&\quad-\omega v_t(x_t,\emptyset;\theta)\end{aligned}\tag{3}

其中,cc\emptyset 分别表示条件和零条件,ω\omega 是控制保真度与多样性平衡的 CFG strength。

2.2 概览

ZipVoice 的架构如 Figure 1 所示,由 text encoder 与 vector field estimator 两部分组成,二者都以 Zipformer 为 backbone。其优势将在 §2.3 说明。

待合成文本先被 tokenized 为 y=(y1,y2,,yN)y=(y_1,y_2,\ldots,y_N),其中 yiy_i 是第 ii 个 text token,NN 是 token 数量。Text encoder 将其变换为 text feature y^RF×N\hat{y}\in\mathbb{R}^{F\times N}FF 为 text-feature dimension。对应的 speech feature 记为 x1RD×Tx_1\in\mathbb{R}^{D\times T},其中 DD 是 feature dimension,TT 是帧长度。之后,对 text feature 执行 Average Upsampling,得到 text condition zRF×Tz\in\mathbb{R}^{F\times T}

为获得 zero-shot TTS 能力,作者沿用 [2] 的 speech infilling 任务。具体地,对 speech feature x1RD×Tx_1\in\mathbb{R}^{D\times T} 应用 binary temporal mask m{0,1}D×Tm\in\{0,1\}^{D\times T},其中 1 表示被 mask 的位置。模型在给定 speech condition (1m)x1(1-m)\odot x_1、插值后的 noisy speech feature xt=(1t)x0+tx1x_t=(1-t)x_0+tx_1 与 text condition zz 时,重建 mx1m\odot x_1。如 Figure 1 所示,三种输入具有相同时间长度,并沿 feature dimension 拼接后送入 vector field estimator。

于是,公式(1)的 CFM objective 可改写为:

LCFM-TTS=Et,q(x1),p0(x0)(vt(xt,z,(1m)x1;θ)(x1x0))m2(4)L_{\text{CFM-TTS}}=E_{t,q(x_1),p_0(x_0)}\left\lVert\left(v_t(x_t,z,(1-m)\odot x_1;\theta)-(x_1-x_0)\right)\odot m\right\rVert^2\tag{4}

这里用 mm mask 训练损失,从而丢弃 speech-condition 位置上的损失。

标准版 ZipVoice 使用上述 Flow Matching 损失训练。作者还提出一个进一步以 flow distillation fine-tune 的加速版本。

训练完成后,模型用 ODE solver 生成 speech feature,再通过额外的 vocoder 变换为波形。作者还设计了改善语音质量的 inference-time 策略,详见 §2.6。

2.3 以 Zipformer 为 Backbone 的 ZipVoice 架构

ZipVoice 包含 Zipformer-based text encoder 与 Zipformer-based vector field estimator,其中大部分参数位于后者。Zipformer 原本是为 automatic speech recognition(ASR)设计的 encoder;与标准 Transformer 相比,它有三项关键设计:U-Net-like 架构、convolution module,以及 attention-weight reuse。

Zipformer 适合作为 vector field estimator backbone,原因有三。第一,U-Net 会在不同分辨率上处理 feature representation,被普遍认为是 diffusion model 中有效的 inductive bias [18], [19]。第二,convolutional neural network(CNN)擅长捕捉细粒度局部模式,可补充 Transformer 对长程全局依赖的建模能力;语音相邻 hidden state 高度相关 [20], [21],因此 CNN 尤其适合增强 TTS 建模。

第三,Zipformer 在同一层的三个模块之间复用 attention weight:两个 self-attention module 与一个 non-linear attention(NLA)module。共享 query/key projection 降低了参数量,而 attention-weight sharing 也提升了计算效率。因此,尽管 Zipformer 原为 ASR 设计,这些特性使它很适合充当 vector field estimator backbone。

近期一些 Flow Matching 模型 [3], [4] 直接把 text embedding 输入 vector field estimator,省去专用 text encoder;作者则发现,设计良好的 text encoder 能提升语音可懂度。ZipVoice 的 Zipformer-based text encoder 保留 convolution module 与 attention-weight reuse,但去掉 U-Net 结构,与已有 hybrid CNN-Transformer 设计 [22], [23] 一致。

2.4 使用 Average Upsampling 的语音—文本对齐

NAR-TTS 的核心挑战之一,是处理 text feature 与 speech feature 的长度差异,即语音—文本对齐。传统 NAR-TTS 在训练时需要显式对齐,例如基于 ASR 的 forced alignment [2] 或 monotonic alignment search [22];推理时还需 duration predictor。这不仅让训练复杂化,错误的时长估计还可能损害语音自然度 [7]。

近期一些 NAR-TTS [3], [5], [24], [25] 不再依赖显式对齐。代表性方法 E2-TTS [3] 用 filler token 把文本补到语音长度,使模型隐式学习对齐,从而显著简化 NAR-TTS 架构。

但这一简单方案存在对齐不准、收敛缓慢的问题 [3]。F5-TTS 加入 ConvNeXt [26] module 来 refine 补齐后的 text condition;ZipVoice 则采用无参数的 Average Upsampling,并作出一个直观假设:同一句话中的所有 token 时长相同。当文本有 NN 个 token、语音有 TT 帧时,每个 token 的时长计算为:

d=TN(5)d=\left\lfloor\frac{T}{N}\right\rfloor\tag{5}

floor operation \lfloor\cdot\rfloor 保证扩展后的 text feature 不超过 speech-feature length。在实践中通常成立的 TNT\geq N 条件下,最短 token duration 为 1。

每个 text embedding 重复 dd 次,text-feature length 从 NN 扩展到 dNd\cdot N。若 T>dNT>d\cdot N,再补入 TdNT-d\cdot N 个 filler embedding。最终得到的 zRF×Tz\in\mathbb{R}^{F\times T} 用作 vector field estimator 的 text condition。

虽然 uniform-duration 假设在理论上过于简单,与真实时长差距不小,但它能为 Flow Matching 模型提供合理的初始 text condition。实验上,这个直接的策略显著提高了对齐准确度。

2.5 用 Flow Distillation 加速 ZipVoice

本节介绍 ZipVoice-Distill:通过 flow distillation 从 ZipVoice 得到的更快变体。作者让 teacher model 执行两步推理,每一步都使用 CFG,据此构造 teacher vector field;再让 student prediction 回归这一 vector field。这样,student 可用更少 NFE 接近 teacher 的表现,并且无需额外的 CFG forward。

给定预训练 TTS teacher θT\theta^T,用其参数初始化 student θS\theta^S。为了让 student 获得 CFG 推理收益却避免 CFG 的额外 model evaluation,作者让 student 显式以 CFG strength 为条件 [27]:先用 Fourier embedding 与 linear layer 处理 ω\omega,再像注入 timestep 一样把它整合进 vector field estimator。

对于任意 timestep tt 与输入 noisy speech xtx_t,teacher 连续走两步,到达中间时刻 tmidt_{\text{mid}} 和目标时刻 tdestt_{\text{dest}}

xtmid=Φ(xt,t,tmid,c,ω;θT)xtdest=Φ(xtmid,tmid,tdest,c,ω;θT)(6)\begin{aligned}x_{t_{\text{mid}}}&=\Phi(x_t,t,t_{\text{mid}},c,\omega;\theta^T)\\x_{t_{\text{dest}}}&=\Phi(x_{t_{\text{mid}}},t_{\text{mid}},t_{\text{dest}},c,\omega;\theta^T)\end{aligned}\tag{6}

其中 one-step ODE solver Φ\Phi 定义为:

Φ(xt,t,tmid,c,ω;θT)=xt+(tmidt)v~t(xt,c,ω;θT)(7)\Phi(x_t,t,t_{\text{mid}},c,\omega;\theta^T)=x_t+(t_{\text{mid}}-t)\tilde{v}_t(x_t,c,\omega;\theta^T)\tag{7}

这里 v~t(xt,c,ω;θT)\tilde{v}_t(x_t,c,\omega;\theta^T) 是按公式(3)使用 CFG 后的 prediction。

两段 step size,即 tmidtt_{\text{mid}}-ttdesttmidt_{\text{dest}}-t_{\text{mid}},都从 [0,Δtmax][0,\Delta t_{\text{max}}] 均匀采样;ω\omega 也从给定区间 [ωmin,ωmax][\omega_{\text{min}},\omega_{\text{max}}] 均匀采样。作者发现,动态采样这些值优于使用固定值。

随后,teacher vector field 计算为:

vT=xtdestxttdestt(8)v^T=\frac{x_{t_{\text{dest}}}-x_t}{t_{\text{dest}}-t}\tag{8}

最后,flow distillation loss 为:

LFD=Et,q(x1),p0(x0)(vt(xt,e^,(1m)x1;θ)vT)m2(9)L_{\text{FD}}=E_{t,q(x_1),p_0(x_0)}\left\lVert\left(v_t(x_t,\hat{e},(1-m)\odot x_1;\theta)-v^T\right)\odot m\right\rVert^2\tag{9}

由于公式(6)的 teacher prediction 使用了 CFG,flow distillation 后,student 只需把 CFG strength ω\omega 作为 TTS 输入之一,就能获得 CFG 收益,而不必在每一步执行双倍 model evaluation。

使用固定 teacher θT\theta^T 完成第一阶段 distillation 后,可从最新 student θS\theta^S 出发进行第二阶段。区别在于 teacher vector field 的构造:此时采用 student 的 exponential moving average(EMA)版本 θ~\tilde{\theta},更新规则为:

θ~S=(1β)θS+βθ~(10)\tilde{\theta}^S=(1-\beta)\theta^S+\beta\tilde{\theta}\tag{10}

该式在每次训练更新后执行,β\beta 是 EMA decay factor。由持续演化的 student 派生 teacher vector field,使第二阶段 distillation 能迭代地 refine student 表现。

2.6 推理策略

作为 zero-shot TTS,ZipVoice 除待合成文本 ysynthesisy^{\text{synthesis}} 外,还需要 audio prompt sprompts^{\text{prompt}} 及其 transcription cpromptc^{\text{prompt}},以模仿参考声音。

合成语音的句子时长,根据 prompt transcription 与待合成文本的 token-length ratio 估计:

Tsynthesis=Tpromptysynthesisyprompt(11)T^{\text{synthesis}}=T^{\text{prompt}}\cdot\frac{|y^{\text{synthesis}}|}{|y^{\text{prompt}}|}\tag{11}

其中,TpromptT^{\text{prompt}} 是 prompt audio 的句子时长。

Text encoder 的输入是 tokenized ysynthesisy^{\text{synthesis}}yprompty^{\text{prompt}} 的拼接;随后用 Average-Upsampled text feature 构造 text condition。Audio condition 则把 audio prompt 补齐到 Tprompt+TsynthesisT^{\text{prompt}}+T^{\text{synthesis}};初始 noisy speech 从标准 Gaussian distribution 采样;最后用 ODE solver 生成合成语音。

作者采用 time-dependent CFG,在说话人相似度与可懂度之间取得更好平衡:早期 NFE 的 unconditional prediction 只丢弃 text condition;后期则同时丢弃 text 与 audio condition。

3 相关工作

加速 TTS 一直是关键挑战。早期 end-to-end TTS [28], [29] 因 AR 语音生成而推理缓慢;FastSpeech [20] 以 NAR 架构改善速度,但早期 NAR 建模能力有限,常生成模糊或 over-smoothed 的语音 [30]。Diffusion model [31] 可改善质量,却依赖大量 NFE。此后,多项工作以 Flow Matching [12] 降低 NFE [2], [32];另一些工作通过 consistency distillation、consistency training [33] 或 ReFlow [34],配合第二阶段训练或辅助目标减少 NFE [35]–[39]。与这些方向并行,也有方法使用高效结构加速 TTS [32], [40]。

不过,多数相关工作聚焦小规模数据,无法在多项指标上匹配 SOTA。ZipVoice 在保持 SOTA 级综合表现的同时取得显著加速;而且它专门研究没有显式语音—文本对齐的 NAR-TTS 加速,这是更困难且研究不足的问题。

4 实验设置

4.1 数据集

为与现有 SOTA zero-shot TTS 比较,作者在 10 万小时 Emilia [10] 上训练 ZipVoice;同时在 585 小时 LibriTTS [8] 上训练小规模模型,用于开发与 ablation。评测覆盖三个常用 benchmark:LibriSpeech-PC [41] 的 test-clean subset [4],含 1127 个英文样本;Seed-TTS test-en [7],含 1088 个来自 Common Voice [42] 的英文样本;以及 Seed-TTS test-zh,含 2020 个来自 DiDiSpeech [43] 的中文样本。

4.2 模型

ZipVoice 由 Zipformer-based text encoder 与 Zipformer-based vector field estimator 组成。Text encoder 有 4 层 Zipformer,每层 encoder dimension 为 192、feedforward dimension 为 512。Vector field estimator 含 5 个 Zipformer encoder stack,downsampling rate 依次为 [1×,2×,4×,2×,1×][1\times,2\times,4\times,2\times,1\times],层数依次为 [2,2,4,4,4][2,2,4,4,4];每层 embedding dimension 为 512、feedforward dimension 为 1536。ZipVoice 总参数量约 123M。

4.3 训练

在 Emilia / LibriTTS 上,ZipVoice 先分别用 Flow Matching objective 训练 1M / 60k updates,再用 flow distillation objective 训练 62k / 12k updates;总 batch size 分别为 4k / 2k 秒。Flow Matching 阶段以 20% 概率丢弃 text condition,用于 CFG。Speech infilling 的 mask length 从 speech-feature length 的 70%–100% 随机采样。Emilia 使用 phoneme token,LibriTTS 使用 character token。

4.4 推理

采样使用 Euler ODE solver。生成的 speech feature 由在 LibriTTS 上训练的预训练 Vocos [44] vocoder 转为波形。

4.5 指标

作者使用三项可复现的 model-based 指标。可懂度以合成语音 transcription 与输入文本之间的 word error rate(WER)衡量:Seed-TTS test-en 使用 Whisper-large-v3 [45],Seed-TTS test-zh 使用 Paraformer-zh [46],LibriSpeech-PC test-clean 使用基于 HuBERT 的 ASR [47]。说话人相似度使用基于 WavLM [48] 的 ECAPA-TDNN [49] 提取 speaker embedding,并计算原始 prompt 与合成语音的 cosine similarity,记为 SIM-o [2]。自然度使用神经网络 MOS predictor UTMOS [50]。

主观评测使用 Comparative Mean Opinion Score(CMOS)与 Similarity Mean Opinion Score(SMOS),分别衡量合成语音的相对质量与说话人相似度。

5 实验结果

5.1 大规模数据上的 Baseline 对比

Table I 将 ZipVoice 与在大规模数据上训练的 SOTA zero-shot TTS 比较,包括三个 AR 模型——CosyVoice [51]、CosyVoice 2 [52]、Spark-TTS [53]——以及三个 NAR 模型——MaskGCT [5]、E2-TTS [3]、F5-TTS [4]。AR 结果取自 [53];NAR 中,E2-TTS 使用 [4] 的非官方实现,其他模型使用官方实现。作者还用 F5-TTS 官方代码训练了一个更小版本,以展示缩小模型造成的性能下降。

Table 1:大规模数据上训练的 zero-shot TTS 模型评测结果。粗体为最佳结果。* 表示结果来自论文;§ 表示非官方实现;† 表示使用官方 checkpoint 推理;‡ 表示作者使用官方代码自行训练。
模型数据(小时)参数量客观指标主观指标
LibriSpeech-PC test-cleanSeed-TTS test-enSeed-TTS test-zh
SIM-o ↑WER ↓UTMOS ↑SIM-o ↑WER ↓UTMOS ↑SIM-o ↑WER ↓UTMOS ↑CMOS ↑SMOS ↑
真实语音--0.6901.874.100.7342.143.520.7551.252.7803.36
AR 模型
CosyVoice*170K Multi.416M---0.6094.29-0.7233.63---
CosyVoice 2*167K Multi.618M---0.6522.57-0.7481.45---
Spark-TTS*102K Multi.507M---0.5841.98-0.6721.20---
NAR 模型
MaskGCT†100K Emilia1048M0.6912.263.910.7132.883.550.7732.402.63-0.084.10
E2-TTS§ (32 NFE)100K Emilia333M0.7002.493.470.7062.323.210.7131.912.26--
F5-TTS† (32 NFE)100K Emilia336M0.6551.893.890.6641.853.720.7501.532.93-0.033.76
F5-TTS‡ (32 NFE)100K Emilia155M0.6152.103.840.6281.963.660.7331.572.93--
ZipVoice (16 NFE)100K Emilia123M0.6681.643.980.6971.703.820.7511.403.150.173.94
ZipVoice-Distill (8 NFE)100K Emilia123M0.6471.544.110.6701.623.910.7401.343.180.163.88
ZipVoice-Distill (4 NFE)100K Emilia123M0.6571.514.050.6791.643.910.7481.393.160.053.84

尽管模型紧凑,ZipVoice 的表现仍可与参数量大得多的 SOTA 模型相当。它在 WER 与 UTMOS 上具有明显优势,SIM-o 也保持可比水平。虽然并非每项指标都最佳,但综合竞争力很强。

对比 ZipVoice 与 ZipVoice-Distill 可见,flow distillation 使后者更偏向语音质量(WER 与 UTMOS),代价是说话人相似度(SIM-o)轻微下降。

Table 2:模型推理速度比较。RTF 越低越好。
模型参数量RTF ↓
GPUCPU
F5-TTS (32 NFE)336M0.295837.284
ZipVoice (16 NFE)123M0.05579.5529
ZipVoice-Distill (8 NFE)123M0.02332.4177
ZipVoice-Distill (4 NFE)123M0.01251.2202

由于 Table 1 中 F5-TTS 最快,Table 2 在相同设备与 PyTorch 版本下比较 F5-TTS 和 ZipVoice 的 real-time factor(RTF)。所有模型都使用同一个 Vocos [44] vocoder,其 RTF 相对 TTS 模型可忽略。实验用 3 秒 prompt 生成 10 秒语音。ZipVoice-Distill(4 NFE)在 NVIDIA H20 GPU 上比 F5-TTS 快 23.7 倍,在 Intel Xeon Platinum 8457C 单 CPU 线程上快 32.6 倍。其速度优势来自三个方面:模型紧凑、NFE 更少、避免 CFG 额外 forward。它在单 CPU 线程上接近实时,可能扩展 SOTA zero-shot TTS 的可部署设备范围。

5.2 小规模数据上的 Baseline 对比

作者还在小规模 LibriTTS 上训练模型,并与同样使用小规模数据的系统比较。Baseline 包括 VALL-E [1]、YourTTS [54] 与 F5-TTS。VALL-E 使用 [55] 的非官方实现,YourTTS 使用官方实现;F5-TTS 则由作者用官方代码在 LibriTTS 上训练 500k updates。

Table 3:小规模数据上训练的 zero-shot TTS 模型评测结果。粗体为最佳。§ 表示非官方实现;† 表示官方 checkpoint 推理;‡ 表示作者使用官方代码自行训练。
模型数据(小时)参数量LibriSpeech-PC test-clean
SIM-o ↑WER ↓UTMOS ↑
真实语音0.6901.874.10
VALL-E§6K Libri-light367M0.3696.203.01
YourTTS†474 Multi.87M0.4626.373.68
F5-TTS (32 NFE)‡555 LibriTTS158M0.5841.784.14
ZipVoice (8 NFE)555 LibriTTS123M0.6101.694.16
ZipVoice-Distill (4 NFE)555 LibriTTS123M0.6061.684.22

如 Table 3 所示,在小规模数据设置下,ZipVoice 仍有竞争力。三个 baseline 中,基于 Flow Matching 的 F5-TTS 最强,说明 Flow Matching model 具有稳健性;但 ZipVoice 与 ZipVoice-Distill 参数更少、训练 updates 更少,仍在所有指标上持续优于 F5-TTS。

5.3 ZipVoice 结构 Ablation

Table 4:ZipVoice 架构 ablation。
模型架构LibriSpeech-PC test-clean
SIM-o ↑WER ↓UTMOS ↑
ZipVoice0.6101.694.16
去掉 text encoder0.5972.044.15
去掉 Average Upsampling0.51320.193.89
+ 使用 ConvNeXt0.52415.493.80

本节检验 ZipVoice 的两项设计:text encoder 与 Average Upsampling。Table 4 表明,两者都对保持高可懂度(低 WER)很重要;尤其去掉 Average Upsampling 后,性能大幅下降。作者也测试了 [4] 的 ConvNeXt-based text-condition refinement;虽然它能一定程度降低 WER,但 Average Upsampling 的优势依然明显。

5.4 Zipformer Backbone Ablation

Table 5:Backbone 结构 ablation。
Backbone 结构LibriSpeech-PC test-clean
SIM-o ↑WER ↓UTMOS ↑
Zipformer0.6101.694.16
去掉 convolution0.5869.794.01
去掉 downsampling0.5573.704.02
去掉 downsampling 与 bypass0.5626.354.08
去掉 bypass0.17998.891.25
去掉 NLA0.5481.833.99
不共享 attention weight0.5951.754.18

作者通过详细 ablation 验证 Zipformer 作为 vector field estimator backbone 的有效性。Table 5 显示,convolution module 对高可懂度很重要;U-Net-like downsampling 与 bypass 在所有指标上都持续有益,说明这种 inductive bias 对 Flow Matching TTS 有效。只去掉 downsampling、保留 bypass 时,架构近似 [3] 的 flat U-Net-style linked Transformer,结果优于同时去掉 downsampling 与 bypass;反之,保留 downsampling 却去掉 bypass 会导致严重退化,突出了 cross-resolution bypass 的重要性。Attention-weight reuse 方面,在 NLA 中复用权重可持续改善全部指标;两个 self-attention module 共享 attention weight 的结果与 baseline 相近,但效率更高。

5.5 不同 Distillation 方法

Table 6:不同 NFE 下的 distillation 方法 ablation。粗体表示每个 NFE 设置中的最佳结果。
Distillation 方法NFELibriSpeech-PC test-clean
SIM-o ↑WER ↓UTMOS ↑
未蒸馏模型10.17192.171.30
20.47515.001.83
40.6342.113.84
Consistency distillation10.44120.721.53
20.5714.063.08
40.5681.973.30
ReFlow10.51217.401.90
20.6015.363.39
40.6082.563.92
本文 flow distillation10.39818.811.58
20.5882.333.73
40.6061.684.22

作者把所提 flow distillation 与两种常用 Flow Matching 加速方法比较:consistency distillation [33] 和 ReFlow [34],且所有模型训练相同 update 数。Table 6 显示,consistency distillation 在 1、2 NFE 时有效,但到 4 NFE 不再改善。由于 text condition 没有显式 token-level duration,ZipVoice 很难在仅 1、2 NFE 时达到满意表现,因此 consistency distillation 不适合该场景。ReFlow [34] 在所有 NFE 下都提高 UTMOS,却在 NFE 大于 1 时恶化 WER。相比之下,本文 flow distillation 在不同 NFE 下都持续改善性能,并在 4 NFE 时取得最佳综合结果。

6 结论

本文提出 ZipVoice:一种快速、高质量、基于 Flow Matching 的 zero-shot TTS。Vector field estimator 使用 Zipformer backbone,在保持强建模能力的同时提高参数效率;Average Upsampling 与 Zipformer text encoder 共同保证稳定的语音—文本对齐与可懂度;flow distillation 则减少 NFE,显著加速推理。实验表明,ZipVoice 以更少参数和更快推理达到可与 SOTA zero-shot TTS 匹配的语音质量。

参考文献

参考文献保留原始书目信息与顺序;正文编号可悬浮、键盘聚焦或轻触查看引用卡。

  1. S. Chen, C. Wang, Y. Wu, Z. Zhang, L. Zhou, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li et al., "Neural codec language models are zero-shot text to speech synthesizers," IEEE Transactions on Audio, Speech and Language Processing, 2025.
  2. M. Le, A. Vyas, B. Shi, B. Karrer, L. Sari, R. Moritz, M. Williamson, V. Manohar, Y. Adi, J. Mahadeokar et al., "Voicebox: Text-guided multilingual universal speech generation at scale," Advances in neural information processing systems, vol. 36, pp. 14 005–14 034, 2023.
  3. S. E. Eskimez, X. Wang, M. Thakker, C. Li, C.-H. Tsai, Z. Xiao, H. Yang, Z. Zhu, M. Tang, X. Tan et al., "E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts," in 2024 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2024, pp. 682–689.
  4. Y. Chen, Z. Niu, Z. Ma, K. Deng, C. Wang, J. Zhao, K. Yu, and X. Chen, "F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching," arXiv preprint arXiv:2410.06885, 2024.
  5. Y. Wang, H. Zhan, L. Liu, R. Zeng, H. Guo, J. Zheng, Q. Zhang, X. Zhang, S. Zhang, and Z. Wu, "MaskGCT: Zero-shot text-to-speech with masked generative codec transformer," in The Thirteenth International Conference on Learning Representations, 2025.
  6. H.-H. Guo, Y. Hu, K. Liu, F.-Y. Shen, X. Tang, Y.-C. Wu, F.-L. Xie, K. Xie, and K.-T. Xu, "Fireredtts: A foundation text-to-speech framework for industry-level generative speech applications," arXiv preprint arXiv:2409.03283, 2024.
  7. P. Anastassiou, J. Chen, J. Chen, Y. Chen, Z. Chen, Z. Chen, J. Cong, L. Deng, C. Ding, L. Gao et al., "Seed-tts: A family of high-quality versatile speech generation models," arXiv preprint arXiv:2406.02430, 2024.
  8. H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, "Libritts: A corpus derived from librispeech for text-to-speech," in Proc. Interspeech 2019, 2019, pp. 1526–1530.
  9. W. Kang, X. Yang, Z. Yao, F. Kuang, Y. Yang, L. Guo, L. Lin, and D. Povey, "Libriheavy: A 50,000 hours asr corpus with punctuation casing and context," in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp. 10 991–10 995.
  10. H. He, Z. Shang, C. Wang, X. Li, Y. Gu, H. Hua, L. Liu, C. Yang, J. Li, P. Shi et al., "Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation," in 2024 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2024, pp. 885–890.
  11. K. Shen, Z. Ju, X. Tan, E. Liu, Y. Leng, L. He, T. Qin, S. Zhao, and J. Bian, "Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers," in ICLR, 2024.
  12. Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le, "Flow matching for generative modeling," in The Eleventh International Conference on Learning Representations, 2023.
  13. X. Li, Z. Shang, H. Hua, P. Shi, C. Yang, L. Wang, and P. Zhang, "Sf-speech: Straightened flow for zero-shot voice clone," IEEE Transactions on Audio, Speech and Language Processing, 2025.
  14. S. Kim, K. Shih, J. F. Santos, E. Bakhturina, M. Desta, R. Valle, S. Yoon, B. Catanzaro et al., "P-flow: A fast and data-efficient zero-shot tts through speech prompting," Advances in Neural Information Processing Systems, vol. 36, pp. 74 213–74 228, 2023.
  15. A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, "Attention is all you need," Advances in neural information processing systems, vol. 30, 2017.
  16. Z. Yao, L. Guo, X. Yang, W. Kang, F. Kuang, Y. Yang, Z. Jin, L. Lin, and D. Povey, "Zipformer: A faster and better encoder for automatic speech recognition," in The Twelfth International Conference on Learning Representations, 2024.
  17. J. Ho and T. Salimans, "Classifier-free diffusion guidance," in NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021.
  18. C. Si, Z. Huang, Y. Jiang, and Z. Liu, "Freeu: Free lunch in diffusion u-net," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 4733–4743.
  19. Y. Tian, Z. Tu, H. Chen, J. Hu, C. Xu, and Y. Wang, "U-dits: Downsample tokens in u-shaped diffusion transformers," in The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024.
  20. Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, "Fastspeech: Fast, robust and controllable text to speech," Advances in neural information processing systems, vol. 32, 2019.
  21. A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu et al., "Conformer: Convolution-augmented transformer for speech recognition," in Proc. Interspeech 2020, 2020, pp. 5036–5040.
  22. J. Kim, S. Kim, J. Kong, and S. Yoon, "Glow-tts: A generative flow for text-to-speech via monotonic alignment search," Advances in Neural Information Processing Systems, vol. 33, pp. 8067–8077, 2020.
  23. C. Miao, S. Liang, M. Chen, J. Ma, S. Wang, and J. Xiao, "Flow-tts: A non-autoregressive network for text to speech based on flow," in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, pp. 7209–7213.
  24. D. Yang, D. Wang, H. Guo, X. Chen, X. Wu, and H. Meng, "Simplespeech: Towards simple and efficient text-to-speech with scalar latent transformer diffusion models," Proc. INTERSPEECH, 2024.
  25. K. Lee, D. W. Kim, J. Kim, S. Chung, and J. Cho, "DiTTo-TTS: Diffusion transformers for scalable text-to-speech without domain-specific factors," in The Thirteenth International Conference on Learning Representations, 2025.
  26. S. Woo, S. Debnath, R. Hu, X. Chen, Z. Liu, I. S. Kweon, and S. Xie, "Convnext v2: Co-designing and scaling convnets with masked autoencoders," in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 16 133–16 142.
  27. C. Meng, R. Rombach, R. Gao, D. Kingma, S. Ermon, J. Ho, and T. Salimans, "On distillation of guided diffusion models," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 14 297–14 306.
  28. Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio et al., "Tacotron: Towards end-to-end speech synthesis," in Proc. Interspeech 2017, 2017, pp. 4006–4010.
  29. J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al., "Natural tts synthesis by conditioning wavenet on mel spectrogram predictions," in 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2018, pp. 4779–4783.
  30. Y. Ren, X. Tan, T. Qin, Z. Zhao, and T.-Y. Liu, "Revisiting over-smoothness in text to speech," arXiv preprint arXiv:2202.13066, 2022.
  31. V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, and M. Kudinov, "Grad-tts: A diffusion probabilistic model for text-to-speech," in International Conference on Machine Learning. PMLR, 2021, pp. 8599–8608.
  32. S. Mehta, R. Tu, J. Beskow, É. Székely, and G. E. Henter, "Matcha-tts: A fast tts architecture with conditional flow matching," in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp. 11 341–11 345.
  33. Y. Song, P. Dhariwal, M. Chen, and I. Sutskever, "Consistency models," in International Conference on Machine Learning. PMLR, 2023, pp. 32 211–32 252.
  34. X. Liu, C. Gong et al., "Flow straight and fast: Learning to generate and transfer data with rectified flow," in The Eleventh International Conference on Learning Representations, 2023.
  35. Z. Ye, W. Xue, X. Tan, J. Chen, Q. Liu, and Y. Guo, "Comospeech: One-step speech and singing voice synthesis via consistency model," in Proceedings of the 31st ACM International Conference on Multimedia, 2023, pp. 1831–1839.
  36. W. Guan, Q. Su, H. Zhou, S. Miao, X. Xie, L. Li, and Q. Hong, "Reflow-tts: A rectified flow model for high-fidelity text-to-speech," in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp. 10 501–10 505.
  37. Y. Guo, C. Du, Z. Ma, X. Chen, and K. Yu, "Voiceflow: Efficient text-to-speech with rectified flow matching," in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp. 11 121–11 125.
  38. Z. Ye, Z. Ju, H. Liu, X. Tan, J. Chen, Y. Lu, P. Sun, J. Pan, W. Bian, S. He et al., "Flashspeech: Efficient zero-shot speech synthesis," in Proceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 6998–7007.
  39. K. Wang, W. Guan, S. Lu, J. Yao, L. Li, and Q. Hong, "Slimspeech: Lightweight and efficient text-to-speech with slim rectified flow," in ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, pp. 1–5.
  40. R. Luo, X. Tan, R. Wang, T. Qin, J. Li, S. Zhao, E. Chen, and T.-Y. Liu, "Lightspeech: Lightweight and fast text to speech with neural architecture search," in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021, pp. 5699–5703.
  41. A. Meister, M. Novikov, N. Karpov, E. Bakhturina, V. Lavrukhin, and B. Ginsburg, "Librispeech-pc: Benchmark for evaluation of punctuation and capitalization capabilities of end-to-end asr models," in 2023 IEEE automatic speech recognition and understanding workshop (ASRU). IEEE, 2023, pp. 1–7.
  42. R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber, "Common voice: A massively-multilingual speech corpus," in Proceedings of the Twelfth Language Resources and Evaluation Conference, 2020, pp. 4218–4222.
  43. T. Guo, C. Wen, D. Jiang, N. Luo, R. Zhang, S. Zhao, W. Li, C. Gong, W. Zou, K. Han et al., "Didispeech: A large scale mandarin speech corpus," in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021, pp. 6968–6972.
  44. H. Siuzdak, "Vocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis," in The Twelfth International Conference on Learning Representations, 2024.
  45. A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, "Robust speech recognition via large-scale weak supervision," in International conference on machine learning. PMLR, 2023, pp. 28 492–28 518.
  46. Z. Gao, S. Zhang, I. McLoughlin, and Z. Yan, "Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition," in Proc. Interspeech 2022, 2022, pp. 2063–2067.
  47. W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, "Hubert: Self-supervised speech representation learning by masked prediction of hidden units," IEEE/ACM transactions on audio, speech, and language processing, vol. 29, pp. 3451–3460, 2021.
  48. S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al., "Wavlm: Large-scale self-supervised pre-training for full stack speech processing," IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1505–1518, 2022.
  49. B. Desplanques, J. Thienpondt, and K. Demuynck, "Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification," in Proc. Interspeech 2020, 2020, pp. 3830–3834.
  50. T. Saeki, D. Xin, W. Nakata, T. Koriyama, S. Takamichi, and H. Saruwatari, "Utmos: Utokyo-sarulab system for voicemos challenge 2022," Interspeech 2022, 2022.
  51. Z. Du, Q. Chen, S. Zhang, K. Hu, H. Lu, Y. Yang, H. Hu, S. Zheng, Y. Gu, Z. Ma et al., "Cosyvoice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic tokens," arXiv preprint arXiv:2407.05407, 2024.
  52. Z. Du, Y. Wang, Q. Chen, X. Shi, X. Lv, T. Zhao, Z. Gao, Y. Yang, C. Gao, H. Wang et al., "Cosyvoice 2: Scalable streaming speech synthesis with large language models," arXiv preprint arXiv:2412.10117, 2024.
  53. X. Wang, M. Jiang, Z. Ma, Z. Zhang, S. Liu, L. Li, Z. Liang, Q. Zheng, R. Wang, X. Feng et al., "Spark-tts: An efficient llm-based text-to-speech model with single-stream decoupled speech tokens," arXiv preprint arXiv:2503.01710, 2025.
  54. E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. Gölge, and M. A. Ponti, "Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone," in International conference on machine learning. PMLR, 2022, pp. 2709–2720.
  55. X. Zhang, L. Xue, Y. Gu, Y. Wang, J. Li, H. He, C. Wang, S. Liu, X. Chen, J. Zhang et al., "Amphion: an open-source audio, music, and speech generation toolkit," in 2024 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2024, pp. 879–884.

LLM WIKI · CONTEXT READER

AI 论文解读

DeepSeek V4 Flash

Enter 发送 · Shift + Enter 换行 · Esc 关闭