带有细粒度控制信号的 text-to-audio(TTA)生成,例如精确的时间控制或可懂语音内容,已在近期工作中得到探索。然而,受数据稀缺限制,这些方法在大规模场景下的生成性能仍然受损。本文将可控 TTA 生成重新表述为一个 multi-task learning 问题,并提出 progressive diffusion modeling 方法 ControlAudio。该方法通过逐步推进的策略,拟合以更细粒度信息为条件的分布,包括文本、时间与音素特征。首先,我们提出一套同时覆盖标注与模拟的数据构建方法,按文本、时间、音素的顺序扩充条件信息。其次,在模型训练阶段,我们先用大规模 text-audio 对预训练 Diffusion Transformer(DiT),获得可扩展的 TTA 生成能力,再利用统一语义表示逐步并入时间与音素特征,扩展模型的可控性。最后,在推理阶段,我们提出 progressively guided generation,依次强化更细粒度的信息,与 DiT 从粗到细的采样特性天然一致。大量实验表明,ControlAudio 在时间准确率和语音清晰度方面达到 state-of-the-art,并在客观与主观评估上显著优于既有方法。演示样例见 https://control-audio.github.io/Control-Audio/。
1 引言
Text-to-audio(TTA)生成系统旨在合成与给定自然语言描述一致的高保真音频,例如“一只鸟正在鸣叫” (Liu et al., 2023;Ghosal et al., 2023;Huang et al., 2023;Evans et al., 2024)。近期研究开始探索对 TTA 系统更细粒度的控制,大体可分为两类。第一类加入精确时间控制,例如“一只鸟在 2–5 秒鸣叫”,相关创新涵盖条件建模技术 (Wang et al., 2025;Xie et al., 2024) 与 training-free latent manipulation (Jiang et al., 2025)。第二类研究可懂音频生成,例如“一只鸟在鸣叫,同时一个男人说:‘今天天气非常晴朗’”,其做法是引入额外模块,同时编码一般音频与语音语义信息 (Lee et al., 2024;Jung et al., 2025)。但是,带有精确时间和语音信息的大规模 text-audio 数据集采集成本很高,这些方法的大规模可控生成性能仍然有限;此前也没有工作在统一框架中探索“时间可控且语音可懂的 TTA 生成”,例如“一只鸟在 0–5 秒鸣叫,随后一个男人在 7–10 秒说:‘今天天气非常晴朗’”。
本文提出 ControlAudio,一种 progressive diffusion modeling 方法,用于逐步捕获以细粒度信息 为条件的目标分布,从而实现可扩展的可控 TTA 生成。我们的设计覆盖数据构建与表示、模型训练以及 guided sampling;每个环节都逐步并入更细粒度的条件信息,由此在规模化条件下扩展可控性。在数据构建方面,我们先收集大规模 对,再通过标注和模拟方法构建成本更高的 与 数据集,为各训练阶段预先定义目标分布。对于文本与时间信息,我们设计 Structured Prompt,使预训练文本编码器无需 fine-tuning 便能准确编码二者。给定时间指示,也就是语音事件的持续时长后,我们进一步用音素 token 自然扩展同一个编码器的词表,以单一文本编码器统一建模文本、时间与音素特征。
在上述数据集预先定义各阶段目标分布后,我们引入 progressive diffusion training:预训练阶段先实现高质量 TTA 合成,持续学习阶段再逐步纳入细粒度控制信号。第一阶段在直接由波形空间压缩得到的 latent space 中预训练 DiT,仅以文本指示为条件,从而获得大规模高保真 TTA。第二阶段同时以文本和时间为条件 fine-tune latent DiT,使模型能够精确控制每个声音事件的时间窗口。可控 TTA 常见的问题是:加入细粒度控制后,无这些控制条件时的 text-conditioned 合成质量会下降 (Wang et al., 2025)。因此,ControlAudio 在第二阶段交替使用纯文本条件和 条件,以避免 progressive training 中的 catastrophic forgetting。最后阶段建立在前两阶段学得的音频生成先验之上,在 text、 与 三种条件之间切换并继续训练 diffusion model,从而在灵活指示下合成高保真音频。
生成时,diffusion model 呈现从粗到细的采样特性:在轨迹早期生成大尺度特征,随后合成细粒度细节,并通过迭代不断完善结果。可控 TTA 的条件信号同样具有不同控制粒度。因此,对时间可控且语音可懂的音频生成,我们设计 Progressively Guided Sampling:先由时间条件引导采样,把各事件的时间窗口作为大尺度特征建立起来;再引入音素条件,把语音内容作为小尺度特征细化。与固定 guidance signal 相比,我们的方法逐步强化更细粒度的条件信息,与 diffusion sampling 过程天然对齐。大量实验表明,ControlAudio 在可控音频生成任务上达到 state-of-the-art,并在时间准确率和语音清晰度的客观、主观评估中显著优于现有方法。
2 相关工作
2.1 可控 TTA 生成
近期工作主要通过两种范式探索 TTA 的时间控制。MC-Diffusion (Guo et al., 2024)、PicoAudio (Xie et al., 2024) 与 Audio ControlNet (Zhu et al., 2026) 等 training-based 方法依赖预定义事件类别,因此对 open-domain prompt 的灵活性受限。较新的 AudioComposer (Wang et al., 2025)、PicoAudio2 (Zheng et al., 2025) 和 DegDiT (Liu et al., 2026) 以自然语言作为控制接口,但在描述复杂多事件场景时经常存在歧义。另一类 TG-Diff (Du et al., 2024) 与 FreeAudio (Jiang et al., 2025) 等 training-free 方法在推理时强制时间对齐,却通常具有较高计算开销,并且难以处理事件密集的场景。与此同时,生成可懂语音仍是一项基础性难题:多数 TTA 模型生成的语音只是含混的人声。VoiceLDM (Lee et al., 2024)、VoiceDiT (Jung et al., 2025) 与 UmbraTTS (Glazer et al., 2025) 虽能生成高质量语音,但它们是专门的 TTS 系统,缺少对一般音频事件的统一控制。CoVoMix2 (Zhang et al., 2025) 等可控对话工作也主要关注纯语音场景。为克服这些限制,我们提出一个统一框架,同时支持一般音频事件的时间可控生成和可懂的多说话人对话。
2.2 Progressive Modeling
在以多种控制信号为条件的视频或 avatar generation 等近期 cross-modal generation 任务中 (Zhao et al., 2025;Zhu et al., 2026;Lin et al., 2025;Hu et al., 2025),progressive modeling 已被证明能够有效处理多条件视频生成。然而,它在 TTA 中的优势只得到局部探索,例如 long-form narrative audio (Guo et al., 2025) 和 video-to-audio 生成 (Tian et al., 2025;Dai et al., 2026);尚未扩展到精确时间控制与可懂语音仍未解决的可控 TTA。为填补这一空白,ControlAudio 使用 progressive modeling,在不同层级实现分阶段的可控性与细粒度调节。
3 预备知识
3.1 基于 Diffusion 的 TTA 生成
基于 diffusion 的 TTA 模型 (Peebles & Xie, 2023;Li et al., 2024) 通常学习 data-to-noise forward process (Ho et al., 2020) 的条件化逆过程:从一个随机初始状态出发,以文本 prompt 为条件,经过多个 diffusion step 逐渐去除噪声。该框架由三个主要模块构成:1)audio variational autoencoder(VAE),在保证重建质量的同时把音频样本变换为压缩 latent representation;2)预训练文本编码器,把文本 prompt 编码成条件 embedding;3)latent diffusion model,以文本 embedding 为条件预测去噪后的 audio latent。ControlAudio 采用基于 DiT 的架构以保证 scalability (Evans et al., 2025),并以文本、时间与音素 embedding 为条件,直接生成从波形压缩得到的 latent audio representation,无需 cascaded decoding (Liu et al., 2023;Guo et al., 2024;Dai et al., 2025)。
3.2 Classifier-Free Guidance
Classifier-Free Guidance(CFG)(Ho & Salimans, 2022;Wang et al., 2025) 在采样过程中强化条件信号 的引导作用。在每个采样 step,CFG-guided diffusion model 产生两个预测:条件估计 和无条件估计 。最终预测由 guidance scale 对二者外推得到:
通常,更大的 guidance scale 会让结果与条件更强地对齐,可能提升 fidelity,但会牺牲 diversity。
4 ControlAudio
4.1 动机
如上所述,latent diffusion model 已推动 TTA 生成质量进步,但精确时间控制或可懂语音控制等可控生成的质量仍然有限。虽然已经出现多种创新,大规模合成质量仍受数据稀缺影响。此外,过去研究很少实现真正通用的 TTA:也就是在加入额外细粒度控制信号的同时,仍保留仅以文本为条件时的高保真音频生成能力。
为解决这些问题,我们提出覆盖数据构建与表示、模型训练和 guided sampling 的 progressive diffusion modeling 设计,用单个 diffusion model 实现 text-guided、timing-indicated 且 intelligible 的音频生成。Figure 1 展示了整体 progressive strategy。
4.2 数据集构建
数据稀缺。 TTA 可以收集到数百万规模的公开弱标注 text-audio 对(见 Appendix A.1),足以支撑大规模高质量合成。然而,这些数据通常只有高层文本描述,缺少可控合成所需的细粒度标注 (Wang et al., 2025)。具体而言,训练时间可控且语音可懂的 TTA,需要同时包含一般音频事件、语音和精确时间标注的数据集。此类数据非常少:现有时间标注音频数据集规模有限,且缺少语音片段转写;公开语音数据集又没有可靠时间标签。为克服这一限制,我们首先构建多来源数据集。
标注数据。 数据标注管线从 AudioSet-SL (Hershey et al., 2021) 开始。该数据集具有可靠时间标注,但没有对应语音转写。构建 ControlAudio 数据集时,我们先筛出所有包含 “human speech” 的 clip,再使用受 MTV (Weng et al., 2025) 启发的 dual-demixing strategy,结合 MVSEP (Solovyev) 与 Spleeter (Hennequin et al., 2020),从每个 clip 提取干净语音轨。随后利用原始时间戳把语音轨切分成独立事件。最后使用 Gemini 2.5 Pro〔注 1〕转写每个分段事件。管线与 prompt 设计的更多细节见 Appendix A.3 和 Appendix G。转写把条件扩展为 ,从而支持细粒度控制。例如,通用标注(man speaking, <3.00,5.00>)会变成包含具体内容的事件(man speaking: “It's been raining all day.”, <3.00,5.00>)。
模拟数据。 为进一步扩充数据集,我们依据真实世界的数据分布构建大规模模拟数据。首先分析 AudioSet-SL 中的语音活动模式,得到统计先验,详见 Appendix A.4。这些分布指导两类场景按比例合成:单说话人场景(monologue)由 LibriTTS-R 中同一说话人的多条 utterance 组成;多说话人场景(dialogue)则从不同说话人采样。组合语音样本后,我们为 utterance 模拟合理的时间排列。最后,将合成的语音与 WavCaps (Mei et al., 2024) 和 VGG-Sound (Chen et al., 2020) 中的非语音背景混合,信噪比从 2–10 dB 的均匀分布采样 (Jung et al., 2025)。该管线额外生成了 171,246 个复杂音频场景,显著提升训练数据规模与多样性。
4.3 统一语义建模
为编码文本、时间与音素等多样条件信息,我们提出 unified semantic modeling 方法:用单个文本编码器以 progressive、coarse-to-fine 的方式处理全部条件。该方法先建立稳健的结构化表示,避免引入多个专用模块的复杂性 (Lee et al., 2024),为渲染音频中的细粒度内容提供简洁而有效的方案。
用于文本和时间表示的 Structured Prompt。 我们方法的基础是 Structured Prompt(),这是一种用于明确、无歧义地定义声学场景组成的新表示。它使用 special token 构成标准格式,把事件描述及其精确起止时间分隔开,如 Figure 2 所示。该格式旨在克服自由自然语言控制的关键缺陷。自然语言往往有歧义:例如 “an alarm sounds from low to high from 1 second to 9 seconds” 中,模型必须判断 “from…to” 指音高变化还是时间边界。随着场景变复杂,自然语言描述还会越来越冗长、难以解析。相比之下,Structured Prompt 简洁、可扩展且 machine-readable,为生成复杂、时间对齐的音频提供稳健基础。
用于音素表示的 Structured Prompt。 可懂语音合成建立在 prompt 提供的时间结构上。关键洞见是:分配给每个语音事件的显式时间窗口(<start,end>)天然定义了 utterance 的总时长。该设计与近期 TTS 研究一致 (Anastassiou et al., 2024;Lee et al., 2024;Chen et al., 2025);这些研究强调 duration information 是引导语音生成的有效约束,可以显式建模,也可以条件化于整句时长。在此基础上,我们把时间信息进一步组织成结构化、事件级时间窗口,实现更细粒度、更可解释的生成控制。
这一简化使得用同一个文本编码器逐步建模粗粒度时间结构与细粒度语音内容既自然又高效。因此,我们在 phoneme level 表示语音内容,例如 “hello” [HH, AH0, L, OW1]。与 word 相比,phoneme 是更直接、面向发音的信号,可减少歧义并改善生成语音的声学一致性。将这些 phoneme token 加入单一编码器的词表后,编码器便学会在指定时间边界内渲染准确的音素序列,并自然获得处理语音时长的能力。
Chain-of-Thought(CoT)LLM Planning。 给定用户的自由形式描述,我们用基于 CoT 的 LLM 将其转换为 Structured Prompt。模型把输入分解为时间对齐的音频事件,在适用时推断语音内容,再把它们组织成带显式时间与音素级信息的统一表示。该过程消除自然语言歧义,并确保语义内容与时间结构一致对齐,从而为可控音频生成提供可靠控制信号。
4.4 Progressive Model Training
为了训练多条件音频生成模型,我们采用 progressive 三阶段训练策略。模型由此逐步获得细粒度控制能力,每个新阶段都建立在此前技能之上并继续细化,保证学习过程稳定、高效。每一阶段均采用 conditional diffusion objective (Ho et al., 2020) 优化,即训练网络预测加入 clean audio latent 的噪声 :
其中, 是 timestep 的 noisy latent, 是 denoising DiT, 是条件信号, 是文本编码器。progressive strategy 的核心在于训练各阶段如何组织与使用条件信号 。
Stage 1:TTA Pre-training。 首先在大规模 text-audio 数据集上预训练 DiT (Evans et al., 2025),学习从文本描述到 audio latent representation 的稳健通用映射,保证高保真的 text-guided audio generation。
Stage 2:Timing-Controlled TTA Fine-tuning。 随后在具有精确时间标注的音频数据集上 fine-tune 预训练模型,同时继续保留无时间信息的纯文本条件训练。该阶段专门优化模型解析同时包含文本与时间的 Structured Prompt,实现 text-guided 且 timing-controlled 的音频生成。
Stage 3:Timing-Controlled and Intelligible TTA Joint Training。 最后解冻文本编码器,对时间控制与语音可懂度进行 joint optimization。模型使用完整多来源数据集训练,其中混合了带时间标注的真实音频和大规模模拟数据。最终阶段使模型能够以连贯、真实的方式联合生成时间可控的一般音频与语音,解决 text-guided、timing-controlled 且 intelligible 的音频生成问题。
总体而言,progressive model training 在继承此前基础能力的同时逐步获得更细粒度的能力。值得注意的是,Stage 3 的 joint optimization 不仅解锁语音可懂度,还进一步提升此前学得的时间精度。我们把显著提升归因于两个因素。第一,带时间标注的语音数据提供了更丰富、更有针对性的信号,用于学习语言内容与时间边界的对齐。第二,fine-tune 文本编码器,使其能与 diffusion backbone 联合优化;这种协同训练让条件模块(文本编码器)和生成模块(DiT)共同适配复杂的多目标任务。整个过程仍保持在一个简洁有效的框架中:单个文本编码器负责处理文本、时间与音素等全部条件信号。
| 方法 | 时间控制(客观) | 生成质量(客观) | 主观 | 效率 | ||||
|---|---|---|---|---|---|---|---|---|
| Eb↑ | At↑ | FAD↓ | KL↓ | CLAP↑ | Temporal↑ | OVL↑ | RTF↓ | |
| 真实音频 | 43.37 | 67.53 | - | - | 0.377 | 4.52 | 4.48 | - |
| AudioLDM Large | 6.79 | 35.66 | 3.95 | 2.46 | 0.260 | 1.84 | 2.40 | 1.141 |
| AudioLDM 2 Large | 7.75 | 42.41 | 3.07 | 1.92 | 0.279 | - | - | 1.496 |
| AudioLDM 2 Full Large | 6.93 | 20.47 | 3.68 | 2.15 | 0.283 | - | - | 1.496 |
| Tango | 1.60 | 26.51 | 2.82 | 1.93 | 0.245 | 1.68 | 2.58 | 1.207 |
| Stable Audio * | 11.28 | 51.67 | 1.93 | 1.75 | 0.318 | 1.94 | 3.44 | 0.821 |
| CCTA | 14.57 | 18.27 | - | - | - | - | - | 1.207 |
| MC-Diffusion | 29.07 | 47.11 | - | - | - | - | - | - |
| Tango + LControl | 21.46 | 55.15 | - | - | - | - | - | 1.207 |
| AudioComposer-Small | 43.51 | 60.83 | 4.92 | 2.00 | 0.261 | 3.12 | 2.52 | 0.721 |
| AudioComposer-Large | 44.40 | 63.30 | - | - | - | - | - | - |
| TG-Diff | 26.70 | 60.06 | 2.66 | - | 0.244 | - | - | 1.207 |
| FreeAudio | 44.34 | 68.50 | 1.92 | 1.73 | 0.321 | - | - | 1.166 |
| ControlAudio | 55.58 | 79.52 | 2.61 | 1.85 | 0.325 | 4.17 | 3.41 | 0.821 |
| ControlAudio full | 49.85 | 71.55 | 1.47 | 1.30 | 0.356 | 3.96 | 3.75 | 0.821 |
4.5 Progressively Guided Sampling
为同时处理时间信息和更细粒度的音素内容,我们提出 Progressively Guided Sampling。该方法以阈值 timestep 为界,把 reverse diffusion process 分成两个阶段,并相应调整 conditioning prompt 与 guidance scale。在初始采样阶段 ,使用不含音素内容的简化 Structured Prompt 和较低 guidance scale 引导模型,促使模型先为全部音频事件建立合理的时间结构:
在剩余采样过程 中,切换到包含完整音素的 Structured Prompt ,并使用更高的 guidance scale 。
第二阶段严格要求结果遵循音素序列,确保在既有结构内合成高可懂度语音。该 coarse-to-fine 策略将事件放置与内容渲染解耦,从而提高时间准确率和语音清晰度。
5 实验
5.1 实验设置
实现细节。 模型采用以 DiT 为 backbone 的 latent diffusion 框架,在 DAC-VAE (Evans et al., 2024) 学得的压缩 latent space 中生成音频。autoencoder 以 16 kHz 采样率运行,latent 时间分辨率为 25 Hz。diffusion model 遵循 Stable Audio (Evans et al., 2024;Evans et al., 2025) 的 DiT 架构,在 latent space 中生成 10 秒音频片段;通过 cross-attention 连接预训练 Flan-T5 large 文本编码器 (Chung et al., 2024),文本、时间和音素信息先统一为 Structured Prompt,再编码进共享 embedding space。我们先预训练 TTA diffusion model 1M steps,然后进行两个优化阶段:时间可控生成 0.5M steps,以及解冻文本编码器后的 joint training 0.5M steps。每一阶段从前一阶段 checkpoint 初始化,形成连续的 multi-stage training pipeline。所有阶段均在 8 张 NVIDIA A800 GPU 上训练,总 batch size 为 128。模型架构与训练详情见 Appendix B,训练和推理伪代码见 Appendix D。
评估数据集。 为客观评估各项能力,我们使用多个成熟数据集。时间可控生成采用 AudioCondition (Guo et al., 2024) 的公开 test split,其细粒度时间标注适合该任务;可懂语音生成采用 AC-Filtered (Lee et al., 2024)。为与先前方法可比,我们按相应格式重写 prompt,示例见 Appendix F.2。一般 TTA 性能在 AudioCaps test set (Kim et al., 2019) 上报告;特定消融还使用 LibriTTS-R 与 LibriSpeech (Panayotov et al., 2015) 的 test-clean split。
评估指标。 评估覆盖时间控制、音频质量和语音可懂度。时间控制遵循既有工作 (Guo et al., 2024;Wang et al., 2025),报告由 sound event detection(SED)系统 (Mesaros et al., 2016) 计算的 event-based measure(Eb)与 clip-level macro F1(At)。音频质量使用 Fréchet Audio Distance(FAD)、Kullback–Leibler(KL)divergence、Fréchet Distance(FD)、Inception Score(IS)(Liu et al., 2023) 与 CLAP (Wu et al., 2023)。语音可懂度同时做客观与主观测试:客观测试用 Whisper Large-v3 (Radford et al., 2023) 转写生成语音并计算 WER;主观测试由 20 名参与者按五分制评价 Speech Intelligibility、Overall Quality(OVL)与和 prompt 的 Relevance(REL)。细节见 Appendix C。
| 方法 | 客观 | 主观 | |||||||
|---|---|---|---|---|---|---|---|---|---|
| FAD↓ | KL↓ | FD↓ | IS↑ | CLAP↑ | WER↓ | Intelligible↑ | OVL↑ | REL↑ | |
| 真实音频 | - | - | - | - | 0.523 | 17.47 | 4.16 | 4.45 | 4.50 |
| AudioLDM 2 Speech | 23.55 | 3.58 | 102.84 | 1.52 | 0.078 | 32.74 | 2.85 | 1.92 | 1.60 |
| VoiceLDM-S | 4.46 | 1.52 | 47.08 | 3.40 | 0.479 | 43.21 | 2.62 | 2.55 | 2.51 |
| VoiceLDM-M | 5.90 | 1.43 | 46.40 | 3.16 | 0.458 | 8.84 | 4.18 | 3.64 | 3.47 |
| VoiceDiT | 4.60 | - | - | - | 0.220 | 7.09 | - | - | - |
| ControlAudio | 3.52 | 1.45 | 32.55 | 4.43 | 0.513 | 6.84 | 4.31 | 4.15 | 3.82 |
5.2 主要结果
时间可控音频生成。 我们将 ControlAudio 与多种 state-of-the-art TTA 模型比较,包括 AudioLDM (Liu et al., 2023)、AudioLDM 2 (Liu et al., 2024)、Tango (Ghosal et al., 2023) 以及内部实现的 Stable Audio (Evans et al., 2024;Evans et al., 2024;Evans et al., 2025)。此外还纳入显式时间条件方法 MC-Diffusion (Guo et al., 2024)、AudioComposer (Wang et al., 2025),以及 training-free baseline TG-Diff (Du et al., 2024) 与 FreeAudio (Jiang et al., 2025),覆盖多类可控 TTA 范式。TG-Diff 在 training-free 框架下报告时间与音质指标,但使用与其他 baseline 不同的 SED model (Turpault et al., 2019)。CCTA 是只使用控制条件、不输入文本的 MC-Diffusion 变体;Tango + LControl 则为 Tango 增加基于语言的时间控制。AudioComposer 采用其公开 Small 版本。评估从时间对齐、音质与效率三方面展开;效率以单张 NVIDIA A800 上的 real-time factor(RTF)衡量 (Liu et al., 2024;Liu et al., 2024)。Table 1 表明,ControlAudio 的时间对齐与现有方法相比具有竞争力或更优,同时在客观、主观音质指标上持续改进。AudioCondition 的 Eb/At 分类别明细见 Appendix E.1。重要的是,这些提升没有引入额外推理开销,在可控性、质量与效率之间形成良好折中。
可懂音频生成。 我们进一步在 AC-Filtered 上评估 ControlAudio 生成可懂语音的能力,并与 AudioLDM 2 Speech、VoiceLDM-S、VoiceLDM-M 和 VoiceDiT 等面向语音的 baseline 比较。对原生不支持时间控制的 baseline,先用 LLM 根据 caption 预测合理时间窗口,详见 Appendix F.2。比较中使用 VoiceLDM 的公开 checkpoint,而 VoiceDiT 直接引用原论文结果。Table 2 显示,ControlAudio 相比全部 baseline 获得更低 WER、更好音质以及更强 text-audio 语义对齐。主观评估同样表明,语音可懂度、整体音质和文本相关性均有改善,说明 ControlAudio 在保留一般音频 fidelity 的同时能生成更清晰、更忠实的语音片段。
Text-to-Audio 生成。 为确认时间与语音内容控制不会损害一般 TTA 能力,我们在标准自然语言 caption 条件下用 AudioCaps test set 评估 ControlAudio,并与 text-to-audio 模型和其他可控音频生成 baseline 比较。既有可控方法常以音质换控制精度,而 ControlAudio 在提供细粒度控制的同时维持较高生成性能。Table 3 表明,ControlAudio 在多项音质指标上达到与 state-of-the-art baseline 相当或更好的结果。这说明 Structured Prompt 条件与词表扩展可以无缝接入 T2A 系统,在不降低语义对齐或声学 fidelity 的前提下实现精确时间与可懂语音控制。
| 方法 | FAD↓ | KL↓ | FD↓ | IS↑ | CLAP↑ |
|---|---|---|---|---|---|
| 真实音频 | - | - | - | - | 0.525 |
| AudioGen | 1.82 | 1.69 | - | - | - |
| AudioLDM | 4.96 | 2.17 | 29.29 | 8.13 | 0.373 |
| AudioLDM 2 | 2.12 | 1.54 | 33.18 | 8.29 | 0.281 |
| Tango | 1.73 | 1.27 | 24.42 | 7.70 | 0.315 |
| Tango 2 | 2.63 | 1.12 | 20.66 | 9.09 | 0.375 |
| Stable Audio * | 1.52 | 1.51 | 18.30 | 13.79 | 0.538 |
| AudioComposer-S | 3.63 | 1.76 | 27.57 | - | - |
| AudioComposer-L | 2.52 | 1.39 | 19.25 | - | - |
| VoiceLDM-S | 13.83 | 3.36 | 63.42 | 4.56 | 0.217 |
| VoiceLDM-M | 9.70 | 2.81 | 55.80 | 4.60 | 0.272 |
| VoiceDiT | 3.55 | 1.87 | - | - | 0.450 |
| ControlAudio | 1.56 | 1.31 | 14.20 | 14.49 | 0.535 |
5.3 消融实验
Prompt 设计消融。 为单独评估 Structured Prompt 的有效性,我们将使用普通自然语言描述训练的 baseline 与使用 Structured Prompt 训练的模型比较。关键是,两者都只训练到 progressive curriculum 的 Stage 2,即专门学习时间控制的阶段,因此可以公平隔离 prompt 格式本身的影响。Table 4 显示,Structured Prompt 模型在 AudioCondition 上持续获得更好的时间对齐和整体音质。结果说明,结构化格式在事件与时间跨度之间建立更清晰、无歧义的映射;复杂场景中,自然语言描述越冗长,时间准确率越容易下降,这一优势越明显。
| 格式 | Eb↑ | At↑ | FAD↓ | KL↓ | CLAP↑ |
|---|---|---|---|---|---|
| NL | 46.23 | 65.36 | 4.11 | 2.25 | 0.245 |
| SP | 51.62 | 70.81 | 3.61 | 2.05 | 0.293 |
| NL full | 40.79 | 61.06 | 1.03 | 1.36 | 0.376 |
| SP full | 43.76 | 64.82 | 0.92 | 1.27 | 0.419 |
词表粒度消融。 为确定可懂语音的最佳词表粒度,我们在 LibriTTS-R 与 LibriSpeech test-clean 上比较 word-level、sub-word(BPE)和 phoneme-level 三种模型,并列出强 VoiceLDM baseline 作为参照。指标包括可懂度的 WER 与语音自然度的 UTMOS(UT-M)(Saeki et al., 2022)。LibriSpeech 的 WER 按既有工作 (Shen et al., 2023) 使用基于 HuBERT 的 ASR model (Hsu et al., 2021) 计算。Table 5 表明,phoneme-level 模型持续且显著优于其他粒度,取得最低 WER 与最高 UTMOS。这证明 phoneme 能更直接地表示口语内容,使 prompt 与声学输出对齐得更紧密,最终得到明显更清晰、可懂度更高的语音。
| Token 类型 | LibriTTS-R | LibriSpeech | ||
|---|---|---|---|---|
| WER↓ | UT-M↑ | WER↓ | UT-M↑ | |
| 真实音频 | 3.75 | 4.17 | 2.15 | 4.06 |
| VoiceLDM-S | 36.65 | 2.59 | 38.61 | 2.76 |
| VoiceLDM-M | 4.98 | 2.83 | 9.76 | 2.77 |
| Word | 6.96 | 4.12 | 6.44 | 4.14 |
| BPE | 7.53 | 4.15 | 5.04 | 4.20 |
| Phoneme | 4.00 | 4.18 | 3.62 | 4.22 |
采样策略分析。 我们通过分析验证 progressive sampling。该 coarse-to-fine 方法先使用低 guidance scale 和不含内容的简化 prompt 建立时间结构,再切换到高 scale 和包含完整音素的 prompt 渲染可懂语音。Figure 4 对不同 、 的结果显示出明确折中:较低初始 scale 对整体音质至关重要,较高后续 scale 对语音可懂度必不可少。实验得到最优配置 。在总计 个采样 step 时,转换 timestep 设为 。我们还分析 的敏感性:较小 (更长的低 guidance 阶段)改善音质;较大 会减少音素细化 step,损害可懂度。性能在合理范围内保持稳定,详细结果见 Appendix E.2。
6 结论
本文提出 ControlAudio,把可控 TTA 重新表述为由 progressive diffusion modeling 解决的 multi-task learning 问题。该 progressive 方法贯穿数据构建、模型训练与推理,使模型逐步掌握来自文本、时间与音素条件的细粒度控制。大量实验表明,ControlAudio 在时间准确率与语音清晰度上达到 state-of-the-art。与此同时,本工作可能被滥用于制造欺骗性内容或冒充他人声音,这凸显出研究稳健检测方法与实施负责任 AI 治理的紧迫性。
致谢
本工作得到教育部基础学科与交叉学科突破计划(No. JYB2025XDXM101)以及国家自然科学基金(62550004、U24A20342、U25B6003、92570001)支持。
局限性
尽管结果很有前景,本工作仍有若干局限。第一,ControlAudio 首次在 timing-controlled TTA 框架中生成可懂语音,但控制主要局限于语音内容;当前框架缺少显式操控情感、韵律或说话人身份等关键风格属性的机制。第二,高质量一般音频与可懂语音之间仍存在根本张力。虽然模型统一了两项任务,但在复杂共现场景中,过度优化其中一种模态可能轻微影响另一种模态的 fidelity。最后,模型性能天然受限于大规模、富标注 audio-speech 数据集的可用性,而此类数据仍然稀缺。现有标注数据与模拟数据的组合虽有效,但未来若有更全面、更高质量的训练语料,性能仍有进一步提升空间。
参考文献
参考文献保留原始书目信息与顺序。
- Anastassiou et al. (2024) Philip Anastassiou, Jiawei Chen, Jitong Chen, Yuanzhe Chen, Zhuo Chen, Ziyi Chen, Jian Cong, Lelai Deng, Chuang Ding, Lu Gao, and 1 others. 2024. Seed-tts: A family of high-quality versatile speech generation models. arXiv preprint arXiv:2406.02430.
- Chen et al. (2020) Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman. 2020. Vggsound: A large-scale audio-visual dataset. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 721–725. IEEE.
- Du et al. (2024) Tianjiao Du, Jun Chen, Jiasheng Lu, Qinmei Xu, Huan Liao, Yupeng Chen, and Zhiyong Wu. 2024. Controllable text-to-audio generation with training-free temporal guidance diffusion. In 2024 IEEE International Conference on Multimedia and Expo (ICME), pages 1–6. IEEE.
- Evans et al. (2024a) Zach Evans, CJ Carr, Josiah Taylor, Scott H Hawley, and Jordi Pons. 2024a. Fast timing-conditioned latent audio diffusion. In Forty-first International Conference on Machine Learning.
- Evans et al. (2024b) Zach Evans, Julian D Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons. 2024b. Long-form music generation with latent diffusion. arXiv preprint arXiv:2404.10301.
- Evans et al. (2025) Zach Evans, Julian D Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons. 2025. Stable audio open. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5. IEEE.
- Gemmeke et al. (2017) Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter. 2017. Audio set: An ontology and human-labeled dataset for audio events. In 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 776–780. IEEE.
- Ghosal et al. (2023) Deepanway Ghosal, Navonil Majumder, Ambuj Mehrish, and Soujanya Poria. 2023. Text-to-audio generation using instruction guided latent diffusion model. In Proceedings of the 31st ACM International Conference on Multimedia, pages 3590–3598.
- Guo et al. (2024) Zhifang Guo, Jianguo Mao, Rui Tao, Long Yan, Kazushige Ouchi, Hong Liu, and Xiangdong Wang. 2024. Audio generation with multiple conditional diffusion model. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 18153–18161.
- Hennequin et al. (2020) Romain Hennequin, Anis Khlif, Felix Voituret, and Manuel Moussallam. 2020. Spleeter: a fast and efficient music source separation tool with pre-trained models. Journal of Open Source Software, 5(50):2154.
- Hershey et al. (2021) Shawn Hershey, Daniel PW Ellis, Eduardo Fonseca, Aren Jansen, Caroline Liu, R Channing Moore, and Manoj Plakal. 2021. The benefit of temporally-strong labels in audio event classification. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 366–370. IEEE.
- Ho et al. (2020) Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851.
- Ho and Salimans (2022) Jonathan Ho and Tim Salimans. 2022. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598.
- Hsu et al. (2021) Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM transactions on audio, speech, and language processing, 29:3451–3460.
- Hu et al. (2025) Teng Hu, Zhentao Yu, Zhengguang Zhou, Sen Liang, Yuan Zhou, Qin Lin, and Qinglin Lu. 2025. Hunyuancustom: A multimodal-driven architecture for customized video generation. arXiv preprint arXiv:2505.04512.
- Huang et al. (2023) Rongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren, Luping Liu, Mingze Li, Zhenhui Ye, Jinglin Liu, Xiang Yin, and Zhou Zhao. 2023. Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models. In International Conference on Machine Learning, pages 13916–13932. PMLR.
- Jiang et al. (2025) Yuxuan Jiang, Zehua Chen, Zeqian Ju, Chang Li, Weibei Dou, and Jun Zhu. 2025. Freeaudio: Training-free timing planning for controllable long-form text-to-audio generation. arXiv preprint arXiv:2507.08557.
- Jung et al. (2025) Jaemin Jung, Junseok Ahn, Chaeyoung Jung, Tan Dat Nguyen, Youngjoon Jang, and Joon Son Chung. 2025. Voicedit: Dual-condition diffusion transformer for environment-aware speech synthesis. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5. IEEE.
- Kim et al. (2019) Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim. 2019. Audiocaps: Generating captions for audios in the wild. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 119–132.
- Koizumi et al. (2023) Yuma Koizumi, Heiga Zen, Shigeki Karita, Yifan Ding, Kohei Yatabe, Nobuyuki Morioka, Michiel Bacchiani, Yu Zhang, Wei Han, and Ankur Bapna. 2023. Libritts-r: A restored multi-speaker text-to-speech corpus. arXiv preprint arXiv:2305.18802.
- Lee et al. (2024a) Keon Lee, Dong Won Kim, Jaehyeon Kim, and Jaewoong Cho. 2024a. Ditto-tts: Efficient and scalable zero-shot text-to-speech with diffusion transformer. arXiv e-prints, pages arXiv–2406.
- Lee et al. (2024b) Yeonghyeon Lee, Inmo Yeon, Juhan Nam, and Joon Son Chung. 2024b. Voiceldm: Text-to-speech with environmental context. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 12566–12571. IEEE.
- Li et al. (2024) Chang Li, Ruoyu Wang, Lijuan Liu, Jun Du, Yixuan Sun, Zilu Guo, Zhenrong Zhang, Yuan Jiang, Jianqing Gao, and Feng Ma. 2024. Qa-mdt: Quality-aware masked diffusion transformer for enhanced music generation. arXiv preprint arXiv:2405.15863.
- Lin et al. (2025) Gaojie Lin, Jianwen Jiang, Jiaqi Yang, Zerong Zheng, and Chao Liang. 2025. Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models. arXiv preprint arXiv:2502.01061.
- Liu et al. (2023) Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark D Plumbley. 2023. Audioldm: Text-to-audio generation with latent diffusion models. arXiv preprint arXiv:2301.12503.
- Liu et al. (2024a) Haohe Liu, Yi Yuan, Xubo Liu, Xinhao Mei, Qiuqiang Kong, Qiao Tian, Yuping Wang, Wenwu Wang, Yuxuan Wang, and Mark D Plumbley. 2024a. Audioldm 2: Learning holistic audio generation with self-supervised pretraining. IEEE/ACM Transactions on Audio, Speech, and Language Processing.
- Liu et al. (2024b) Huadai Liu, Rongjie Huang, Yang Liu, Hengyuan Cao, Jialei Wang, Xize Cheng, Siqi Zheng, and Zhou Zhao. 2024b. Audiolcm: Text-to-audio generation with latent consistency models. arXiv preprint arXiv:2406.00356.
- Liu et al. (2024c) Huadai Liu, Jialei Wang, Rongjie Huang, Yang Liu, Heng Lu, Zhou Zhao, and Wei Xue. 2024c. Flashaudio: Rectified flows for fast and high-fidelity text-to-audio generation. arXiv preprint arXiv:2410.12266.
- Mei et al. (2024) Xinhao Mei, Chutong Meng, Haohe Liu, Qiuqiang Kong, Tom Ko, Chengqi Zhao, Mark D Plumbley, Yuexian Zou, and Wenwu Wang. 2024. Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research. IEEE/ACM Transactions on Audio, Speech, and Language Processing.
- Mesaros et al. (2016) Annamaria Mesaros, Toni Heittola, and Tuomas Virtanen. 2016. Metrics for polyphonic sound event detection. Applied Sciences, 6(6):162.
- Panayotov et al. (2015) Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015. Librispeech: an asr corpus based on public domain audio books. In 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 5206–5210. IEEE.
- Peebles and Xie (2023) William Peebles and Saining Xie. 2023. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4195–4205.
- Radford et al. (2023) Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. In International conference on machine learning, pages 28492–28518. PMLR.
- Saeki et al. (2022) Takaaki Saeki, Detai Xin, Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, and Hiroshi Saruwatari. 2022. Utmos: Utokyo-sarulab system for voicemos challenge 2022. Interspeech 2022.
- Shen et al. (2023) Kai Shen, Zeqian Ju, Xu Tan, Yanqing Liu, Yichong Leng, Lei He, Tao Qin, Sheng Zhao, and Jiang Bian. 2023. Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers. arXiv preprint arXiv:2304.09116.
- (36) Roman Solovyev. Cinematic sound demixing. https://github.com/ZFTurbo/MVSEP-CDX23-Cinematic-Sound-Demixing.
- Turpault et al. (2019) Nicolas Turpault, Romain Serizel, Ankit Parag Shah, and Justin Salamon. 2019. Sound event detection in domestic environments with weakly labeled data and soundscape synthesis. In Workshop on Detection and Classification of Acoustic Scenes and Events.
- Wang et al. (2025a) Junyou Wang, Zehua Chen, Binjie Yuan, Kaiwen Zheng, Chang Li, Yuxuan Jiang, and Jun Zhu. 2025a. Audiomog: Guiding audio generation with mixture-of-guidance. arXiv preprint arXiv:2509.23727.
- Wang et al. (2024) Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, and 1 others. 2024. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191.
- Wang et al. (2025b) Yaoting Wang, Shengqiong Wu, Yuecheng Zhang, Shuicheng Yan, Ziwei Liu, Jiebo Luo, and Hao Fei. 2025b. Multimodal chain-of-thought reasoning: A comprehensive survey. arXiv preprint arXiv:2503.12605.
- Wang et al. (2025c) Yuanyuan Wang, Hangting Chen, Dongchao Yang, Zhiyong Wu, and Xixin Wu. 2025c. Audiocomposer: Towards fine-grained audio generation with natural language descriptions. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5. IEEE.
- Weng et al. (2025) Shuchen Weng, Haojie Zheng, Zheng Chang, Si Li, Boxin Shi, and Xinlong Wang. 2025. Audio-sync video generation with multi-stream temporal control. arXiv preprint arXiv:2506.08003.
- Wu et al. (2023) Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2023. Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5. IEEE.
- Xie et al. (2024) Zeyu Xie, Xuenan Xu, Zhizheng Wu, and Mengyue Wu. 2024. Picoaudio: Enabling precise timestamp and frequency controllability of audio events in text-to-audio generation. arXiv preprint arXiv:2407.02869.
- Zhang et al. (2025) Leying Zhang, Yao Qian, Xiaofei Wang, Manthan Thakker, Dongmei Wang, Jianwei Yu, Haibin Wu, Yuxuan Hu, Jinyu Li, Yanmin Qian, and 1 others. 2025. Covomix2: Advancing zero-shot dialogue generation with fully non-autoregressive flow matching. arXiv preprint arXiv:2506.00885.
Appendix A 训练数据集
A.1 预训练数据集与预处理
Table 6 汇总了预训练 TTA backbone 所用的全部语料。为学习稳健的 text-audio 映射,我们聚合多样的大规模公开数据集,其中既有 WavCaps (Mei et al., 2024)、AudioCaps (Kim et al., 2019) 等带描述性 caption 的数据,也有大规模 AudioSet (Gemmeke et al., 2017) 等只带高层事件标签的语料。详细描述与广泛声音类别的结合,使模型能学习稳健、通用的语义表示。
这些来源的全部音频都经过统一预处理。首先重采样至 16 kHz 并转换为单声道。为满足 diffusion model 的定长输入要求,所有 clip 统一处理为 10 秒:不足 10 秒的在右侧补静音,超过 10 秒的随机裁取一个 10 秒片段。
| 数据集 | 时长(h) | 数量 | 文本 |
|---|---|---|---|
| AudioCaps | 109 | 44K | caption |
| WavCaps | 7090 | 400K | caption |
| Clotho v2 | 152 | 7k | caption |
| AudioSet | 5800 | 2M | 类别标签 |
| FSD50k | 108 | 51K | 类别标签 |
| ESC-50 | 2.8 | 2K | 类别标签 |
| VGG-Sound | 550 | 210k | 类别标签 |
| MTT | 200 | 24K | caption |
| MSD | 7333 | 880K | caption |
| FMA | 900 | 11K | caption |
A.2 时间控制数据集
时间控制 fine-tuning 数据集建立在 AudioSet-Strong (Hershey et al., 2021) 之上,该数据集包含 1.8M 条音频 clip,并为 456 类声音事件提供密集的 frame-level 时间戳。AudioSet-Strong 只有 “Dog” 之类类别标签而没有描述性文本,因此我们为每个带时间的事件生成更丰富的 caption。具体做法受 WavCaps (Mei et al., 2024) 启发,使用 LLM 为每个切分后的音频事件创建唯一文本描述。
该数据的预处理与预训练阶段有一个关键差异:为了保持时间戳完整,不能随机裁剪。我们固定取每条 clip 的前 10 秒,再过滤事件标注,只保留时间戳落在 0–10 秒窗口内的事件;不足 10 秒的 clip 在右侧补静音。这一确定性流程保证最终数据集的时间标注与对应音频片段严格对齐。
A.3 标注数据管线
本节逐步说明标注数据管线。它用于构建真实世界语音事件数据,同时提供精确时间边界和文本转写,流程如下。
从 AudioSet-SL 选择初始数据。 AudioSet-SL (Hershey et al., 2021) 为广泛声音事件提供经人工验证的强时间标注。我们从完整数据集中筛选全部 10 秒音频 clip,只要其中至少有一个事件被标为 “Human speech”“Speech” 或其任一子类就保留,最终得到 49,950 条含语音 clip。
提取高质量语音轨。 为获得高质量、可靠的 speech stem,我们采用受 MTV (Weng et al., 2025) 启发的 dual-demixing comparison strategy。该策略比较 MVSEP (Solovyev) 与 Spleeter (Hennequin et al., 2020) 的分离输出,进行质量过滤,并提取供后续处理使用的干净语音信号。
事件级切分。 将干净语音轨切成互不重叠的独立语音事件。切分直接使用 AudioSet-SL 原始人工标注的起止时间戳,每个结果音频段代表原录音中的一条连续语音 utterance。该过程得到 173,831 个独立语音片段,随后送往转写。
使用 Large Language Model 转写。 173,831 个干净、已切分的语音事件逐一送入 Gemini 2.5 Pro,用直接 prompt 生成精确文本转写。为保证最终标注质量,我们明确要求:若片段中的语音不可懂或被噪声严重遮蔽,模型必须返回空输出。该步骤同时充当关键质量过滤器。
从切分到过滤式转写的完整管线最终保留 152,070 个高质量、已转写语音事件。数据集中每个事件都带有精确开始时间、结束时间以及经验证的文本转写,为模型学习真实世界的 timed speech 提供真实且具有挑战性的数据来源。
A.4 模拟数据管线
除真实标注数据外,我们还构建了大规模模拟数据管线。目标是在真实世界数据统计模式指导下,生成具有精确时间与转写信息的真实、复杂音频场景。流程分为两阶段:推导统计先验,以及在先验指导下进行合成。
从 AudioSet-SL 推导统计先验。 为让模拟数据反映真实世界的语音活动模式,我们分析 AudioSet-SL 中的含语音 clip,确定两个关键分布。
- 说话人分布: 约 79.1% 的 clip(49,950 条中的 39,509 条)只有一个说话人,20.9% 有多个说话人。该比例指导模拟数据中 monologue 与 dialogue 场景的比例。
- 每条 clip 的 utterance 数分布: 经验分布在每条 clip 只有一条 utterance()处达到峰值,占全部单说话人场景的 32.20%。当 时,clip 频率通常随 utterance 数增加而下降,形成长尾。模拟 monologue 时从该分布采样 utterance 数,并把每条 clip 的上限设为 8,以聚焦最常见场景。完整分布见 Table 7。
| 事件数() | 数量 | Percentage (%) |
|---|---|---|
| 1 | 12,723 | 32.20 |
| 2 | 6,462 | 16.36 |
| 3 | 6,284 | 15.90 |
| 4 | 5,720 | 14.48 |
| 5 | 4,201 | 10.63 |
| 6 | 2,328 | 5.89 |
| 7 | 1,047 | 2.65 |
| 8 | 456 | 1.15 |
| 9 | 150 | 0.38 |
| 10 | 67 | 0.17 |
| 11 | 41 | 0.10 |
| 12 | 20 | 0.05 |
| 10 | 0.04 | |
| 合计 | 39,509 | 100.00 |
Guided Synthesis Pipeline。 每条 10 秒 clip 的合成步骤如下。首先从说话人分布采样场景类型,单说话人 monologue 的概率为 79.1%。接着从 LibriTTS-R (Koizumi et al., 2023) 获取带转写的干净语音 utterance。对 monologue,从同一说话人采样若干 utterance,其数量来自上述分布且最多为 8;对 dialogue,从 2–4 个不同说话人采样,并保证每位说话人最多贡献 4 条 utterance。随后在 10 秒窗口内为这些 utterance 模拟合理时间排列。最后,把合成的纯语音轨与 WavCaps (Mei et al., 2024)、VGG-Sound (Chen et al., 2020) 过滤子集中随机选择的一条非语音背景混合,SNR 从 2–10 dB 的均匀分布采样。
Appendix B 模型配置
本节说明在 fine-tune 为 ControlAudio 之前,用于预训练的 base model 架构。diffusion model 在 latent diffusion modeling(LDM)范式中采用 DiT(Diffusion Transformer)架构。预训练时,模型接收三类条件:自然语言 prompt、开始时间 seconds_start 和总时长 seconds_total,全部嵌入 768 维 feature space。prompt 由预训练 Flan-T5 large 编码,而 seconds_start 与 seconds_total 作为数值输入处理。
diffusion network backbone 是一个 24 层、24 个 attention head、hidden dimension 为 1536 的 DiT (Evans et al., 2024)。模型对全部条件输入使用 cross-attention,并对 duration-related signal 使用 global conditioning。diffusion model 的内部 token dimension 为 64,conditional token dimension 为 768,global condition embedding dimension 为 1536。
B.1 压缩网络
audio autoencoder 是基于 Descript Audio VAE (Evans et al., 2025) 的 VAE,以 16 kHz 采样率运行。模型在大规模公开数据的音频部分从头训练,以学习紧凑 audio representation。encoder 的 model dimension d_model 为 128,stride 为 [4, 4, 4, 10],总下采样率 640。encoder 把输入波形映射为最终 64 维 latent representation,再由 decoder 重建。输入/输出 channel io_channels 设为 1,以处理单声道音频。网络全程使用 Snake activation,decoder 末端不使用 tanh activation。
B.2 训练细节
为提高收敛稳定性与生成质量,我们采用多项常见训练策略 (Evans et al., 2025)。模型参数使用 Exponential Moving Average(EMA);优化器采用 AdamW,learning rate 为 ,,weight decay 为 。
learning-rate schedule 分为两阶段。前 99% 的训练 iteration 保持初始 learning rate 不变;最后 1% 按 InverseLR 公式衰减:
其中, 是衰减阶段内的 step count,,。该策略先在主训练阶段保持稳定、快速收敛,再以短暂的 decaying learning rate fine-tune。
最终阶段从 Stage 2 checkpoint 初始化模型并解冻 Flan-T5 文本编码器,使其与 diffusion backbone 联合优化;其余优化配置沿用前一阶段。joint training 使文本编码器能够适配复合 prompt:其中同时包含 Structured Prompt 格式、时间 special token 和面向语音的 phoneme-level 扩展词表。最终,模型学到统一表示,把语义描述、精确时间跨度和可懂语音内容等多样输入映射成单个高质量、时间可控的音频输出。
Appendix C 评估
C.1 客观指标
我们从音频质量以及与文本 prompt 的语义对齐两个方面进行全面客观评估。
音频质量。 主要 fidelity 指标是 Fréchet Audio Distance(FAD),它基于 VGGish embedding 衡量生成音频与参考音频的分布差异。为评估声学事件分布的一致性,还报告用 PANNs tagging model 计算的 Kullback–Leibler(KL)divergence。为完整起见并与既有工作比较,另将 Inception Score(IS)与 Fréchet Distance(FD)作为补充指标 (Liu et al., 2023)。
语义对齐。 使用 LAION-CLAP score (Wu et al., 2023) 衡量生成音频与文本 prompt 的对齐。该分数定义为生成音频 与文本 prompt 的 CLAP embedding 之间的 cosine similarity:
更高的 CLAP score 表示共享 embedding space 中的语义对应更好。为保持一致,所有客观指标都使用官方 AudioLDM evaluation toolkit 计算。
C.2 主观评估
主观评估招募 20 名评价者,按五分制 Mean Opinion Score(MOS)为生成音频打分(1–5,越高越好)。评估分为两个独立任务,各有专门标准。
时间可控音频生成。 参与者听取音频 clip 并看到对应 timed prompt,然后从以下两个方面评分。
- Temporal Alignment(Temporal): 衡量对时间戳的遵循准确度。问题是:“音频事件的时间与 prompt 给定的开始、结束时间匹配得有多准确?”
- Overall Quality(OVL): 衡量音频 clip 本身的感知质量。问题是:“忽略 prompt,你认为这段音频的整体质量和真实感如何?”
可懂音频生成。 参与者听取包含语音的音频 clip,并看到该语音应表达的文本,再从以下方面评分。
- Speech Intelligibility(Intelligible): 衡量口语内容的清晰度。问题是:“音频中的口语内容有多清晰、易懂?”
- Overall Quality(OVL): 衡量整个声学场景的质量。问题是:“综合语音与全部背景声音,你认为整体音频质量如何?”
- Relevance(REL): 衡量音频与文本之间的语义对应。问题是:“作为整体,这段生成音频与文本描述匹配得如何?”
Appendix D 训练与推理伪代码
Algorithm 1 与 Algorithm 2 汇总 ControlAudio 的训练和推理过程。训练遵循 progressive 三阶段设计,可控性从文本扩展到时间,再扩展到 phoneme-level 语音内容;训练中随机切换不同条件组合,在引入更细粒度控制信号时保留此前能力。推理使用两阶段 progressive sampling:先以低 guidance 建立时间结构,再引入音素条件并提高 guidance,细化可懂语音。
- 输入:训练数据集
- 输出:训练后的 ControlAudio 模型
- 初始化 diffusion model 与预训练文本编码器 。
- Stage 1:Text-to-Audio Pretraining。 每个训练 step 从 采样 ;采样 、;计算 ;预测 ;以 更新 。
- Stage 2:Timing-Controlled Training。 每个训练 step:从 采样 ;随机选择条件 ;采样 ;计算 ;预测 ;更新 。
- Stage 3:Joint Training with Phoneme。 解冻文本编码器 ;每个训练 step:从 采样 ;随机选择条件 ;采样 ;计算 ;预测 ;更新 。
- 输入:文本 prompt 、timestep 阈值 、guidance scale
- 输出:生成音频样本
- 初始化 。
- 构造 与 。
- 从 到 :计算 ,以 scale 应用 CFG,再计算 。
- 从 到 :计算 ,以 scale 应用 CFG,再计算 。
- 返回生成音频 。
| Eb | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| 方法 | 闹钟 | 搅拌机 | 猫 | 餐具 | 狗 | 牙刷 | 煎炸 | 水声 | 语音 | 吸尘器 | 平均 |
| FreeAudio | 40 | 47 | 14 | 20 | 23 | 78 | 60 | 51 | 39 | 72 | 44.34 |
| ControlAudio | 53.08 | 48.57 | 27.06 | 40.55 | 50.41 | 80.51 | 61.15 | 52.01 | 64.23 | 78.26 | 55.58 |
| At | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| 方法 | 闹钟 | 搅拌机 | 猫 | 餐具 | 狗 | 牙刷 | 煎炸 | 水声 | 语音 | 吸尘器 | 平均 |
| FreeAudio | 73 | 67 | 38 | 42 | 59 | 88 | 85 | 81 | 76 | 76 | 68.50 |
| ControlAudio | 87.22 | 69.39 | 41.25 | 67.48 | 85.50 | 91.67 | 91.38 | 82.69 | 88.65 | 90.00 | 79.52 |
| 100 | 95 | 90 | 85 | 80 | |
|---|---|---|---|---|---|
| FAD | 5.82 | 4.74 | 3.85 | 3.40 | 3.35 |
| WER | 5.98 | 6.61 | 6.83 | 7.12 | 7.67 |
| 90 | 89 | 88 | 87 | 86 | 85 | |
|---|---|---|---|---|---|---|
| FAD | 3.85 | 3.61 | 3.52 | 3.53 | 3.47 | 3.40 |
| WER | 6.83 | 6.85 | 6.84 | 6.86 | 7.01 | 7.12 |
Appendix E 详细结果
E.1 AudioCondition 分类别结果
Table 8 给出 AudioCondition 的分事件类别结果,包括 event-based(Eb)与 clip-level(At)指标。ControlAudio 在大多数类别上持续改进,尤其是语音相关和时间结构复杂的事件。
E.2 转换 timestep 的敏感性分析
音质与语音可懂度之间存在清晰 trade-off:较小 改善 FAD,较大 有利于 WER。围绕选定值的细粒度分析还表明,性能在适中范围内保持稳定。综合这一折中,全部实验统一设置 。
Appendix F LLM Planning
F.1 用于 Prompt Planning 的 Chain-of-Thought
近期进展表明 LLM 具有强大的规划和 cross-modal reasoning 能力 (Wang et al., 2025;Wang et al., 2024)。我们让 LLM 充当 “planner”,把自由形式自然语言 caption 自动转换成供生成模型使用的精确 Structured Prompt 。这一转换受 Chain-of-Thought(CoT)范式启发,按 Figure 3 所示分为三个 reasoning stage。
- 事件与时间规划。 给定输入 caption,LLM 首先识别一组不同音频事件 。对每个事件推断一组时间跨度 ,其中 、 分别是以秒计的开始与结束时间。multi-span 表示能够处理多次出现的事件。
- 语音内容规划。 对被识别为语音的事件 ,LLM 推断符合整体语境的合理 utterance 。该步骤为规划事件加入具体、可懂的语音内容,得到中间 tuple 。
- Prompt Recaption。 最后,LLM 把提取的信息序列化为最终 Structured Prompt 。它从原始 caption 开始,为每个规划事件追加特定格式字符串,包含事件名、相关时间跨度以及推断出的语音内容。
| 输入 | 生成的 Structured Prompt(输出) |
|---|---|
| Caption: 她正在公园里说话。 Text: “早上好!你今天感觉怎么样?” | 中文释义: 公园环境声持续 0.00–10.00 秒;女性语音在 1.50–6.00 秒说 “Good morning! How are you feeling today?” 原始 Structured Prompt: She is talking in the park. @{park ambient sounds. & <0.00,10.00>}@{Female speech, woman speaking. & <1.50,6.00> "Good morning! How are you feeling today?"} |
| Caption: 一个孩子大喊,一个小男孩在多次拍击硬表面的声音中说话。 Text: “Say yeah, baby. Say yeah, baby. Are you over tired?” | 中文释义: 小男孩在 1.50–8.00 秒说话;孩子在 2.00–6.00 秒大喊;硬表面拍击声出现在 2.50–3.00 秒和 5.00–5.50 秒。 原始 Structured Prompt: A child yelling as a young boy talks during several slaps on a hard surface. @{Young boy speaking & <1.50,8.00> "Say yeah, baby. Say yeah, baby. Are you over tired?"} @{Child yelling & <2.00,6.00>} @{slaps on a hard surface & <2.50,3.00> <5.00,5.50>} |
| Caption: 一名女性伴随窸窣声说话,随后另一名女性说话。 Text: “The IT services at the King's University College are proud to announce that we have launched” | 中文释义: 第一段女性语音在 0.50–6.00 秒;窸窣声在 1.00–5.00 秒;第二段女性语音在 6.50–8.00 秒。 原始 Structured Prompt: A female speaking with some rustling followed by another female speaking. @{Female speech, woman speaking & <0.50,6.00> "The IT services at the King's University College are proud to announce that"} @{rustling & <1.00,5.00>} @{Female speech, woman speaking & <6.50,8.00> "we have launched"} |
| Caption: 鸭叫之后一名男子说话,远处有鸟鸣。 Text: “Mama Mama snow mama come over here, baby” | 中文释义: 鸭叫在 0.50–1.50 秒;男性语音在 2.00–7.50 秒;远处鸟鸣出现在 2.50–4.00 秒和 5.50–7.00 秒。 原始 Structured Prompt: A duck quacks followed by a man talking while birds chirp in the distance. @{duck quack & <0.50,1.50>} @{Man speaking & <2.00,7.50> "Mama Mama snow mama come over here, baby"} @{birds chirping in the distance & <2.50,4.00> <5.50,7.00>} |
| Caption: 两名男子说话,伴有响亮的昆虫嗡鸣。 Text: “I've got gloves covered in mid repellent. Still fishing.” | 中文释义: 两段男性语音分别位于 1.00–4.50 秒和 5.00–6.50 秒;昆虫嗡鸣覆盖 0.00–10.00 秒。 原始 Structured Prompt: Two men speaking with loud insects buzzing. @{Man speaking & <1.00,4.50> "I've got gloves covered in mid repellent."} @{Man speaking & <5.00,6.50> "Still fishing."} @{loud insects buzzing & <0.00,10.00>} |
| Caption: 男子说话,水流飞溅,远处隐约播放音乐。 Text: “in the amateur show tonight then tomorrow on Saturday the broadcasters and the other amateur cast will be going out hope to do well there get some good footage hope you enjoy” | 中文释义: 男性语音在 0.50–9.50 秒;水流声与远处微弱音乐都覆盖 0.00–10.00 秒。 原始 Structured Prompt: A man speaking as a stream of water splashes and flows while music faintly plays in the distance. @{Man speaking & <0.50,9.50> "in the amateur show tonight then tomorrow on Saturday the broadcasters and the other amateur cast will be going out hope to do well there get some good footage hope you enjoy"} @{water splashing and flowing & <0.00,10.00>} @{faint music in the distance & <0.00,10.00>} |
| Caption: 人们咯咯笑,一名男子说话。 Text: 无。 | 中文释义: 笑声在 1.00–5.00 秒;男性语音在 2.50–4.50 秒说 “What's so funny?” 原始 Structured Prompt: People are giggling, and a man speaks. @{people giggling & <1.00,5.00>} @{Man speaking & <2.50,4.50> "What's so funny?"} |
| Caption: 无。 Text(原文含粗俗用语): “Some people talk about fucking the heads, but the way I do it, I just put my finger down there and pull it out.” 中文直译: “有些人会说要狠狠干那些头,但我的做法是,我只是把手指伸进去,再把它拉出来。” | 中文释义: planner 将无 caption 的输入改写为“一名男子正在给出说明或解释过程”,并把语音放在 1.00–9.00 秒。 原始 Structured Prompt: A person is giving instructions or explaining a procedure. @{Man speaking & <1.00,9.00> "Some people talk about fucking the heads, but the way I do it, I just put my finger down there and pull it out."} |
F.2 AC-Filtered 上的规划结果
为定性评估 LLM prompt planner,Table 11 展示了 AC-Filtered 样本的若干规划结果。planner 能解析复杂、自由形式的 caption(有或没有配套 speech text),并转换为框架所需的精确 machine-readable Structured Prompt。该能力对复杂多说话人场景尤其关键:planner 可以生成 prompt,把不同 utterance 分配给指定时间的不同说话人。相比之下,VoiceLDM 等面向语音的模型即使得到描述对话的 prompt,也只能把全部语音内容渲染成单一 voice 的一段 utterance。规划并生成真实 dialogue 的能力,是本方法创建逼真声学场景的重要优势。
Appendix G 通过 ALM 进行语音转写
我们使用 Gemini 2.5 Pro 为切分后的语音事件生成文本转写。每个干净音频段直接输入模型;所设计的 prompt 同时承担两项功能:准确转写口语内容,以及作为质量过滤器。具体而言,若片段中的语音不可懂或被噪声严重遮蔽,prompt 要求模型返回空字符串,从而自动丢弃低质量样本。该过程确保只有清晰、有效的音频片段才会变成高质量 audio-text 对。完整 prompt 见 Figure 5。