ABSTRACT / 摘要
文本到音频(TTA)系统因能够依据文本描述合成通用音频而受到关注,但既有系统的生成质量有限、计算成本也很高。本文提出 AudioLDM:它在 latent space 中利用对比语言—音频预训练(CLAP)embedding 学习连续音频表征。预训练 CLAP 使 LDM 能在训练时以音频 embedding 为条件,而在采样时改用文本 embedding。由于只学习音频信号的 latent 表征而不在生成模型中另行建模跨模态关系,AudioLDM 同时提升了生成质量与计算效率。仅用一张 GPU 在 AudioCaps 上训练,AudioLDM 在客观与主观指标上均优于其他开源 TTA 系统;它还是首个以零样本方式支持多种文本引导音频操作(如风格迁移)的 TTA 系统。实现与示例见 https://audioldm.github.io。
1. 引言
按照个性化需求生成声效、音乐或语音,对增强现实、虚拟现实、游戏开发和视频编辑等应用十分重要。传统音频生成依赖信号处理技术;近年来,无条件或由其他模态提供条件的生成模型改变了这一任务。早期研究主要处理类别数很少的 label-to-sound 设置,例如 UrbanSound8K 的十个声音类别。相比之下,自然语言更灵活,能够描述音高、声学环境和时间顺序等细粒度属性。依据自然语言描述生成音频的任务称为 text-to-audio(TTA)生成。
TTA 系统需要生成种类广泛的高维音频。为高效建模,DiffSound 使用学习得到的离散表征;AudioGen 又以波形离散表征上的自回归建模推进了这一方向。受到 StableDiffusion 在连续 latent 表征上使用 LDM 获得高质量图像的启发,本文把此前的 TTA 方法扩展到连续 latent space,而不是继续学习离散表征。考虑到游戏等应用还需要风格迁移等音频操作,作者也探索了此前尚未展示过的多种零样本文本引导音频操作。
既有 TTA 工作通常要求大规模、高质量的音频—文本配对,但这类数据既稀缺,质量与数量也都有限。为利用低质量配对数据,一些方法会预处理文本,却可能丢失声事件之间的关系,例如把 “a dog is barking at the park” 变成 “dog bark park”。相比之下,AudioLDM 的生成模型训练只需要音频数据,绕开了文本预处理,并且实验表明它优于直接使用音频—文本配对训练。
AudioLDM 学习由 mel-spectrogram VAE 编码的 latent 表征,并用 CLAP embedding 为 LDM 提供条件。CLAP 已把音频与文本对齐到同一个 embedding space,因此训练 VAE latent 生成器时可直接从音频本身取得条件,不再需要配对的音频—文本数据。实验进一步表明,仅用音频训练 LDM 甚至优于使用配对数据:AudioLDM 在 AudioCaps 上取得 FD 23.31,明显优于 DiffSound 的 47.68,并能在采样过程中实现零样本音频操作。
- 首次为 TTA 构建连续空间 LDM,并在主观与客观指标上超过既有方法。
- 利用 CLAP embedding,使 LDM 无需语言—音频配对数据即可训练。
- 实验证明,仅使用音频数据也能得到高质量、计算高效的 TTA 系统。
- 无需针对任务微调,即可进行文本引导的音频风格迁移、超分辨率与补绘。
2. 相关工作
近期 TTA 系统通常依据自然语言描述学习离散音频表征,再把表征解码为波形。由于 latent 生成器依赖音频—文本配对数据,DiffSound 与 AudioGen 都提出了缓解配对数据稀缺和低质量问题的方法。
DiffSound 由文本编码器、decoder、VQ-VAE 与 vocoder 构成。它用 mask-based text generation(MBTG)从音频标签生成文本描述,例如把 “dog bark, a man speaking” 表示为带有 mask token 的序列。但这种生成文本仍只包含标签信息,可能限制模型性能。
AudioGen 使用 Transformer decoder 生成直接由波形压缩得到的离散 token。它在十个数据集上训练,并通过按不同信噪比混合音频、拼接处理后的文本描述来增强数据。其预处理会把自然语言压缩为标签,例如把 “a dog is barking at the park” 变成 “dog bark park”,因而丢失细致的空间与时间关系。
Diffusion model 已在图像生成、图像修复、语音与波形生成以及视频生成中取得先进质量,但在高维数据空间中反复迭代会导致推理缓慢。一个解决方案是在较小的 latent space 中运行扩散。对 TTA 而言,波形中的冗余信息会增加建模复杂度并降低速度;DiffSound 虽在 mel-spectrogram 的压缩离散 token 上使用扩散,但生成质量有限,也没有探索音频操作。
3. 文本条件音频生成
3.1 对比语言—音频预训练
受到文本到图像模型使用 CLIP 生成图像先验的启发,AudioLDM 使用对比语言—音频预训练(CLAP)来支持 TTA 生成。
对于音频样本 与文本描述 ,文本编码器 和音频编码器 分别提取 与 ,其中 为 CLAP embedding 维度。作者依据既有 CLAP 研究,以 HTS-AT 构建音频编码器、以 RoBERTa 构建文本编码器,并采用对称 cross-entropy loss 训练。训练细节与语言—音频数据见附录 A。
CLAP 训练完成后,音频样本 会被映射为对齐音频与文本的联合空间中的 embedding 。CLAP 已在零样本音频分类等下游任务中表现出泛化能力,因此对未见过的语言或音频样本,其 embedding 也能提供跨模态信息。
3.2 条件 Latent Diffusion Model
给定文本描述 ,TTA 系统生成音频 。LDM 用模型分布 估计真实条件数据分布 。其中, 是由 mel-spectrogram 压缩得到的音频 latent; 为压缩级别, 为通道数, 为时频维。CLAP 把 与 对齐到同一跨模态空间,因此训练 LDM 时可使用音频 embedding,TTA 生成时则可换成文本 embedding。
Diffusion model 包含两个过程:(i)前向过程按照预定义 noise schedule 把数据分布变为标准 Gaussian;(ii)反向过程依据推理 noise schedule 从噪声逐步生成数据样本。
在前向过程中,每个时间步 的转移概率为:
其中, 是注入噪声, 是 的重参数化, 表示每一步的噪声水平。最终时间步 上, 服从标准各向同性 Gaussian。
为优化模型,AudioLDM 使用重加权的噪声估计目标:
是预训练 CLAP 音频编码器 从波形 得到的 embedding。反向过程从 Gaussian noise 和文本 embedding 出发,按下式逐步生成音频 latent :
反向过程的均值与方差参数化为:
是预测的生成噪声,且 。训练阶段学习在音频跨模态表征 条件下生成 ;TTA 生成时再提供文本 embedding 。因此,借助 CLAP,LDM 训练阶段无需文本监督即可实现 TTA。网络结构详见附录 B。
3.3 条件增强
大规模图文配对帮助图像 diffusion model 捕捉对象与背景之间的细粒度关系,但现有语言—音频数据规模远小于语言—图像数据。AudioGen 会混合两段音频并拼接相应的处理后 caption,形成新的配对数据。AudioLDM 训练 LDM 时只以音频 embedding 为条件,因此可以只增强音频信号,而不必同步增强语言—音频对。具体地,作者对 执行 mixup:
从 Beta distribution 采样。因为 LDM 训练不需要文本信息,所以无需构造对应文本 。混合音频对增加了 LDM 的训练样本对 ,使模型对 CLAP embedding 更稳健;采样时,LDM 再根据未见过的语言描述所产生的 生成对应音频 latent 。
3.4 Classifier-free Guidance
Classifier-free guidance(CFG)在每个采样步引导可控生成。训练时,作者以固定概率(例如 )随机丢弃条件 ,从而同时训练条件模型 与无条件模型 。生成时以文本 embedding 为条件,并使用修正后的噪声估计:
其中 决定 guidance scale。与 AudioGen 相比有两点不同:其一,AudioGen 在 Transformer 自回归模型上使用 CFG,而 AudioLDM 的 LDM 保留了 CFG 的 diffusion formulation;其二,AudioLDM 的 来自未经预处理的自然语言,因此能利用描述中的空间与时间细节,而 AudioGen 的文本预处理会删除这些信息。
3.5 Decoder
卷积 VAE 把 mel-spectrogram 压缩到 ,其中 为 latent space 的压缩级别。VAE 由堆叠卷积模块组成 encoder 与 decoder,训练目标结合重建损失、对抗损失和 Gaussian constraint loss。采样时,decoder 根据 LDM 生成的音频 latent 重建 。作者测试 ,并以兼顾效率和质量的 为默认值;同样使用式(8)的数据增强来保证组合样本的重建质量。最后,HiFi-GAN 把重建 mel-spectrogram 转换为波形 。
4. 文本引导音频操作
风格迁移。 给定源音频 ,按照式(2)的前向过程,在预定义时间步 得到带噪 latent 。将它作为预训练 AudioLDM 反向过程的起点,便可通过浅层反向过程 ,用文本 操作源音频:
控制操作幅度。当 时,源音频信息几乎不被保留,操作接近普通 TTA 生成。Figure 3 展示了 的影响:在 时可看到更大幅度的变化。
补绘与超分辨率。 二者都根据已观测部分 生成缺失内容。AudioLDM 把已观测部分的 latent 注入生成 latent :从 开始,每完成一次式(5)的反向推理步,就按下式修改 :
是修改后的 latent, 是 latent space 的观测 mask, 则由式(2)对 加噪得到。卷积 VAE 大致保留 mel-spectrogram 与 latent 的空间对应关系;因此若时频点 已观测,便令 。式(11)借此在保留真实观测 的同时,依据文本 prompt 生成缺失信息。
5. 实验
训练数据集。 本文使用 AudioSet(AS)、AudioCaps(AC)、Freesound(FS)与 BBC Sound Effect library(SFX)。AS 是当时最大的音频数据集,含 个标签和超过 小时音频;AC 较小,约有 个带文本描述的 clip。AudioSet 与 AudioCaps 多为来自 YouTube 的自然场景音频,质量无法保证,因此作者又抓取覆盖音乐、语音和声效的 Freesound 与 BBC SFX,以补充高质量音频。处理与训练配置见附录 E。
评估数据集。 作者在 AC 与 AS 上评估。AC 每个 audio clip 有五条 caption,随机抽取一条作为文本条件。由于 AC 作者有意移除了音乐相关标签的音频,为覆盖更广的声音分布,本文还随机选择 AS 的 作为评估集;AS 没有文本描述,因此把标签拼接为 prompt,例如 “Speech, hip hop music, and crowd cheering”。
评估方法。 客观指标包括 Fréchet distance(FD)、inception score(IS)和 Kullback–Leibler(KL)divergence。FD 衡量生成样本与目标样本的相似度,IS 同时评价质量与多样性,KL 在配对样本级计算后取平均;三者均基于 PANNs 音频分类器。为与 AudioGen 比较,作者也报告基于 VGGish 的 Fréchet audio distance(FAD),但认为 VGGish 的性能可能不及 PANNs,因此以 FD 为主要指标。主观评估招募六名音频专业人士,分别在 到 分范围内评价整体质量(OVL)和与输入文本的相关性(REL)。
模型。 基线包括 DiffSound 与 AudioGen:DiffSound 在 AS 和 AC 上训练,约 M 参数;AudioGen 在 AS、AC 及另外八个数据集上训练,约 M 参数。由于 AudioGen 未公开实现,本文复用其论文报告的 KL 与 FAD。作者训练了 M 参数的 AudioLDM-S 和 M 参数的 AudioLDM-L;为直接展示方法优势,这两个模型只用 AC。为研究训练数据规模,又训练了使用 AC、AS、Freesound 与 BBC SFX 的 AudioLDM-L-Full。UNet 结构见附录 B。
| 模型 | 文本数据 | 使用 CLAP | 参数量 | 时长(h) | FD ↓ | IS ↑ | KL ↓ | FAD ↓ | OVL ↑ | REL ↑ |
|---|---|---|---|---|---|---|---|---|---|---|
| Ground truth | — | — | — | — | — | — | — | — | 83.61 ± 1.1 | 80.11 ± 1.2 |
| DiffSound† | ✓ | ✗ | 400M | 5420 | 47.68 | 4.01 | 2.52 | 7.75 | 45.00 ± 2.6 | 43.83 ± 2.3 |
| AudioGen† | ✓ | ✗ | 285M | 8067 | — | — | 2.09 | 3.13 | — | — |
| AudioLDM-S-Full-RoBERTa | ✓ | ✗ | 181M | 145 | 32.13 | 4.02 | 3.25 | 5.89 | — | — |
| AudioLDM-S | ✗ | ✓ | 181M | 145 | 29.48 | 6.90 | 1.97 | 2.43 | 63.41 ± 1.4 | 64.83 ± 0.9 |
| AudioLDM-L | ✗ | ✓ | 739M | 145 | 27.12 | 7.51 | 1.86 | 2.08 | 64.30 ± 1.6 | 64.72 ± 1.6 |
| AudioLDM-S-Full | ✗ | ✓ | 181M | 8886 | 23.47 | 7.57 | 1.98 | 2.32 | — | — |
| AudioLDM-L-Full | ✗ | ✓ | 739M | 8886 | 23.31 | 8.13 | 1.59 | 1.96 | 65.91 ± 1.0 | 65.97 ± 1.6 |
5.1 结果
在 AC 测试集上,仅使用 AC 训练的 AudioLDM-S 即以更小模型在客观与主观指标上超过基线;扩大模型为 AudioLDM-L 后结果继续提升;再加入 AS、FS 与 SFX 后,AudioLDM-L-Full 取得最佳质量,FD 达 。尽管 RoBERTa 与 CLAP 的文本编码器结构相同,CLAP 已在预训练时对齐音频与文本,把跨模态关系学习和生成模型训练解耦;相反,AudioLDM-S-Full-RoBERTa 的文本编码器只表示文本,生成模型必须同时学习文本—音频关系和生成过程。CLAP 还允许仅用音频数据训练。
主观评估与客观指标趋势一致。AudioLDM 的 OVL 与 REL 都约为 ,显著超过 DiffSound 的 与 。更大的 AudioLDM 对整体音频质量更有利,扩大训练数据后 OVL 和 REL 都明显提高。Figure 4 显示 AudioLDM 的得分比 DiffSound 更集中在高分区域;随机选择的真实录音 spam case 得分很高,说明评分结果可信。
为评估可能包含音乐的音频,作者又在 AS 子集上比较 DiffSound。三种 AudioLDM 模型呈现与 AC 测试集相同的趋势,并在全部指标上大幅优于 DiffSound。
| 模型 | FD ↓ | IS ↑ | KL ↓ |
|---|---|---|---|
| DiffSound | 50.40 | 4.19 | 3.63 |
| AudioLDM-S | 28.08 | 6.78 | 2.51 |
| AudioLDM-L | 27.51 | 7.18 | 2.49 |
| AudioLDM-L-Full | 24.26 | 7.67 | 2.07 |
条件信息。 由于 LDM 训练使用音频 embedding ,TTA 生成却使用文本 embedding ,一个自然问题是直接用文本 embedding 训练能否更强。公平比较时,作者也采用 AudioGen 的增强方式:按 Section 3.3 混合音频对,并把两条 caption 拼接为条件。Table 3 表明,以 训练明显优于以 训练。
| 模型 | 文本 | 音频 | FD ↓ | IS ↑ | KL ↓ |
|---|---|---|---|---|---|
| AudioLDM-S | ✓ | ✓ | 31.26 | 6.35 | 2.01 |
| AudioLDM-S | ✗ | ✓ | 29.48 | 6.90 | 1.97 |
| AudioLDM-S-Full | ✓ | ✓ | 27.20 | 7.52 | 2.38 |
| AudioLDM-S-Full | ✗ | ✓ | 23.47 | 7.57 | 1.98 |
| AudioLDM-L-Full | ✓ | ✓ | 25.79 | 7.95 | 2.26 |
| AudioLDM-L-Full | ✗ | ✓ | 23.31 | 8.13 | 1.59 |
作者认为主要原因是文本 embedding 不如音频 embedding 能准确表示生成目标。声音具有歧义与复杂性,caption 很难完整准确,不同标注者对同一段音频也可能有不同理解;有些 caption 还高度抽象,例如 BBC SFX 中的 “Boats: Battleships-5.25 conveyor space”,人类也难以想象其声音。相比之下,CLAP 的 直接从音频提取,并与理想的最佳 caption 对齐,因而不必依赖带噪文本标签即可向 LDM 提供强条件。
压缩级别。 作者比较 ,发现压缩级别越高,生成质量越低。不过即便 、把 band mel-spectrogram 的频率轴压到仅 个维度,KL 仍与 AudioGen 相当,全部指标也都优于 DiffSound。
| 模型 | FD ↓ | IS ↑ | KL ↓ | |
|---|---|---|---|---|
| AudioLDM-S | 29.48 | 6.90 | 1.97 | |
| AudioLDM-S | 33.50 | 6.13 | 2.04 | |
| AudioLDM-S | 34.32 | 5.68 | 2.09 |
当 (即直接根据 CLAP embedding 生成 mel-spectrogram)或 时,单张 RTX 3090 难以完成训练,推理也会很慢。 在保持较高生成质量的同时把计算量降到合理水平,因此本文默认采用 。
文本引导音频操作。 超分辨率把采样率从 kHz 提升到 kHz;补绘则删除十秒音频中 到 秒的内容,再生成该区间。鉴于多数音频超分辨率研究只处理语音,作者同时在 AudioCaps 与多说话人语音数据集 VCTK 上实验,以 log-spectral distance(LSD)比较 AudioUNet、NVSR 和 AudioLDM;补绘使用 FAD,并在该任务上建立基线。
| 任务 | 超分辨率 | 补绘 | |
|---|---|---|---|
| 数据集 | AudioCaps | VCTK | AudioCaps |
| Unprocessed | 2.76 | 2.15 | 10.86 |
| Kuleshov et al. (2017) | — | 1.32 | — |
| Liu et al. (2022a) | — | 0.78 | — |
| AudioLDM-S | 1.59 | 1.12 | 2.33 |
| AudioLDM-L | 1.43 | 0.98 | 1.92 |
AudioLDM 在超分辨率上优于强基线 AudioUNet,但不及 NVSR。原因之一是 AudioLDM 在包含重背景噪声的多样化音频上训练,其超分辨率输出也可能出现白噪声或其他非语音事件,降低指标。尽管如此,这些结果为 TTA 系统以零样本方式完成文本引导音频操作打开了方向,并建立了可继续改进的 benchmark。
5.2 消融实验
把 UNet 的 attention 简化为单层 multi-head self-attention 后,各项指标明显下降,说明复杂 attention 更合适。音频分类中常用的平衡采样对 TTA 没有改善。条件增强提升了主观 OVL 与 REL,却没有提升客观指标;作者推测,mixup 生成的数据不完全属于 AudioCaps 分布,使生成分布也与评估集错位。考虑到主观质量提升,作者仍建议使用条件增强。
| 设置 | FD ↓ | IS ↑ | KL ↓ | OVL ↑ | REL ↑ |
|---|---|---|---|---|---|
| AudioLDM-S | 29.48 | 6.90 | 1.97 | 63.41 | 64.83 |
| w. Simple attn | 33.12 | 6.15 | 2.09 | — | — |
| w. Balance samp | 34.05 | 6.21 | 2.16 | — | — |
| w. Cond aug | 31.88 | 6.25 | 2.02 | 64.49 | 65.01 |
DDIM 采样步数。 反向过程的推理步数直接影响 DDPM 生成质量:通常步数增加会同时提升质量和计算量。Table 7 显示更多 DDIM 步数带来更好结果,但在约 步后收益趋于饱和, 步只比 步略好。
| DDIM steps | |||||
|---|---|---|---|---|---|
| FD | 55.84 | 42.84 | 35.71 | 30.17 | 29.48 |
| IS | 4.21 | 5.91 | 6.51 | 6.85 | 6.90 |
| KL | 2.47 | 2.12 | 2.01 | 1.94 | 1.97 |
Guidance scale。 guidance 在条件一致性与样本多样性之间折中。 时 FD 与 KL 最佳,但 FAD 不是最佳;作者推测,FAD 使用的音频分类器不如 FD 的 PANNs,增强对细粒度语言描述的遵循反而可能成为误导信息。为与报告 FAD 的既有工作比较,主结果采用 ,同时完整给出 对 FAD、FD、IS 和 KL 的影响。
案例研究。 附录 I 展示风格迁移(Figure 9–11)、超分辨率(Figure 12)、补绘(Figure 13–14)和 TTA 生成(Figure 15–22)。TTA 示例具体展示对声学环境、材料、声事件、音高、音乐流派与时间顺序的控制能力。
6. 结论
本文提出结合 CLAP 与 LDM 的 TTA 方法 AudioLDM,在生成质量、计算效率与音频操作能力上具有优势。仅用 AudioCaps 和一张 GPU,AudioLDM 即在主客观指标上达到当时 SOTA;同时支持零样本文本引导的音频风格迁移、超分辨率与补绘。
7. 致谢
作者感谢 James King 与 Jinhua Liang 就 latent diffusion model 所作的有益讨论。本研究部分得到 BBC Research and Development、EPSRC 项目 EP/T019751/1 “AI for Sound”,以及 University of Surrey 的 CVSSP/FEPS 博士奖学金支持。为实现开放获取,作者已对由此产生的任何 Author Accepted Manuscript 版本采用 Creative Commons Attribution(CC BY)许可。
参考文献
参考文献保留原始书目信息与论文顺序。
- Andresen, U. A new way in sound synthesis. In Audio Engineering Society. Audio Engineering Society, 1979.
- Chen, K., Du, X., Zhu, B., Ma, Z., Berg-Kirkpatrick, T., and Dubnov, S. HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection. In IEEE International Conference on Acoustics, Speech and Signal Processing, 2022a.
- Chen, N., Zhang, Y., Zen, H., Weiss, R., Norouzi, M., and Chan, W. Wavegrad: Estimating gradients for waveform generation. In International Conference on Learning Representations, 2021.
- Chen, Z., Tan, X., Wang, K., Pan, S., Mandic, D., He, L., and Zhao, S. Infergrad: Improving diffusion models for vocoder by considering inference in training. In IEEE International Conference on Acoustics, Speech and Signal Processing, 2022b.
- Chen, Z., Wu, Y., Leng, Y., Chen, J., Liu, H., Tan, X., Cui, Y., Wang, K., He, L., Zhao, S., Bian, J., and Mandic, D. Resgrad: Residual denoising diffusion probabilistic models for text to speech. arXiv preprint:2212.14518, 2022c.
- Dhariwal, P. and Nichol, A. Diffusion models beat gans on image synthesis. In Conference on Neural Information Processing Systems, 2021.
- Drossos, K., Lipping, S., and Virtanen, T. Clotho: an audio captioning dataset. In IEEE International Conference on Acoustics, Speech and Signal Processing, 2020.
- Engel, J., Hantrakul, L., Gu, C., and Roberts, A. Ddsp: Differentiable digital signal processing. arXiv preprint:2001.04643, 2020.
- Gemmeke, J. F., Ellis, D. P., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M. AudioSet: An ontology and human-labeled dataset for audio events. In IEEE International Conference on Acoustics, Speech and Signal Processing, pp. 776–780. IEEE, 2017.
- Gong, Y., Chung, Y.-A., and Glass, J. PSLA: Improving audio tagging with pretraining, sampling, labeling, and aggregation. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:3292–3306, 2021.
- Hershey, S., Chaudhuri, S., Ellis, D. P., Gemmeke, J. F., Jansen, A., Moore, R. C., Plakal, M., Platt, D., Saurous, R. A., Seybold, B., et al. CNN architectures for large-scale audio classification. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing, pp. 131–135. IEEE, 2017.
- Ho, J. and Salimans, T. Classifier-free diffusion guidance. In NeurIPS Workshop on Deep Generative Models and Downstream Applications, 2021.
- Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. In Conference on Neural Information Processing Systems, 2020.
- Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., and Salimans, T. Imagen video: High definition video generation with diffusion models. arXiv preprint:2210.02303, 2022.
- Isola, P., Zhu, J.-Y., Zhou, T., and Efros, A. A. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1125–1134, 2017.
- Karplus, K. and Strong, A. Digital synthesis of plucked-string and drum timbres. Computer Music Journal, 7(2):43–55, 1983.
- Kilgour, K., Zuluaga, M., Roblek, D., and Sharifi, M. Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms. In INTERSPEECH, pp. 2350–2354, 2019.
- Kim, C. D., Kim, B., Lee, H., and Kim, G. Audiocaps: Generating captions for audios in the wild. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 119–132, 2019.
- Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. arXiv preprint:1412.6980, 2014.
- Kingma, D. P. and Welling, M. Auto-encoding variational bayes. arXiv preprint:1312.6114, 2013.
- Kong, J., Kim, J., and Bae, J. Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis. Advances in Neural Information Processing Systems, 33:17022–17033, 2020a.
- Kong, Q., Cao, Y., Iqbal, T., Wang, Y., Wang, W., and Plumbley, M. D. PANNs: Large-scale pretrained audio neural networks for audio pattern recognition. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 28:2880–2894, 2020b.
- Kong, Q., Cao, Y., Liu, H., Choi, K., and Wang, Y. Decoupling magnitude and phase estimation with deep resunet for music source separation. arXiv preprint:2109.05418, 2021a.
- Kong, Z., Ping, W., Huang, J., Zhao, K., and Catanzaro, B. Diffwave: A versatile diffusion model for audio synthesis. In International Conference on Learning Representations, 2021b.
- Kreuk, F., Synnaeve, G., Polyak, A., Singer, U., Défossez, A., Copet, J., Parikh, D., Taigman, Y., and Adi, Y. Audiogen: Textually guided audio generation. arXiv preprint:2209.15352, 2022.
- Kuleshov, V., Enam, S. Z., and Ermon, S. Audio super resolution using neural networks. arXiv:1708.00853, 2017.
- Lam, M., Wang, J., Huang, R., Su, D., and Yu, D. Bilateral denoising diffusion models. In International Conference on Learning Representations, 2022.
- Lee, S., Kim, H., Shin, C., Tan, X., Liu, C., Meng, Q., Qin, T., Chen, W., Yoon, S., and Liu, T. Priorgrad: Improving conditional denoising diffusion models with data-driven adaptive prior. In International Conference on Learning Representations, 2022.
- Leng, Y., Chen, Z., Guo, J., Liu, H., Chen, J., Tan, X., Mandic, D., He, L., Li, X.-Y., Qin, T., et al. Binauralgrad: A two-stage conditional diffusion probabilistic model for binaural audio synthesis. arXiv preprint:2205.14807, 2022.
- Liu, H., Kong, Q., Tian, Q., Zhao, Y., Wang, D., Huang, C., and Wang, Y. Voicefixer: Toward general speech restoration with neural vocoder. arXiv preprint:2109.13731, 2021a.
- Liu, H., Choi, W., Liu, X., Kong, Q., Tian, Q., and Wang, D. Neural vocoder is all you need for speech super-resolution. arXiv preprint:2203.14941, 2022a.
- Liu, H., Kong, Q., Liu, X., Mei, X., Wang, W., and Plumbley, M. D. Ontology-aware learning and evaluation for audio tagging. arXiv preprint:2211.12195, 2022b.
- Liu, H., Liu, X., Kong, Q., Wang, W., and Plumbley, M. D. Learning the spectrogram temporal resolution for audio classification. arXiv preprint:2210.01719, 2022c.
- Liu, N., Li, S., Du, Y., Torralba, A., and Tenenbaum, J. B. Compositional visual generation with composable diffusion models. In European Conference on Computer Vision, 2022d.
- Liu, X., Iqbal, T., Zhao, J., Huang, Q., Plumbley, M. D., and Wang, W. Conditional sound generation using neural discrete time-frequency representation learning. In IEEE International Workshop on Machine Learning for Signal Processing, pp. 1–6. IEEE, 2021b.
- Liu, X., Liu, H., Kong, Q., Mei, X., Plumbley, M. D., and Wang, W. Simple pooling front-ends for efficient audio classification. arXiv preprint:2210.00943, 2022e.
- Liu, X., Liu, H., Kong, Q., Mei, X., Zhao, J., Huang, Q., Plumbley, M. D., and Wang, W. Separate what you describe: language-queried audio source separation. arXiv preprint:2203.15147, 2022f.
- Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. Roberta: A robustly optimized BERT pretraining approach. arXiv preprint:1907.11692, 2019.
- Nichol, A. and Dhariwal, P. Improved denoising diffusion probabilistic models. In International Conference on Machine Learning, 2021.
- Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint:2112.10741, 2021.
- Oord, A. v. d., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K. WaveNet: A generative model for raw audio. arXiv preprint:1609.03499, 2016.
- Pascual, S., Bhattacharya, G., Yeh, C., Pons, J., and Serrà, J. Full-band general audio synthesis with score-based diffusion. arXiv preprint:2210.14661, 2022.
- Perez, E., Strub, F., De Vries, H., Dumoulin, V., and Courville, A. Film: Visual reasoning with a general conditioning layer. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32, 2018.
- Popov, V., Vovk, I., Gogoryan, V., Sadekova, T., and Kudinov, M. Grad-tts: A diffusion probabilistic model for text-to-speech. In International Conference on Machine Learning, 2021.
- Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, 2021.
- Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 2020.
- Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. Hierarchical text-conditional image generation with clip latents. arXiv preprint:2204.06125, 2022.
- Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10684–10695, 2022.
- Saharia, C., Ho, J., Chan, W., Salimans, T., Fleet, D. J., and Norouzi, M. Image super-resolution via iterative refinement. arXiv preprint:2104.07636, 2021.
- Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., Salimans, T., Ho, J., Fleet, D. J., and Norouzi, M. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint:2205.11487, 2022.
- Salamon, J., Jacoby, C., and Bello, J. P. A dataset and taxonomy for urban sound research. In Proceedings of the 22nd ACM international conference on Multimedia, pp. 1041–1044, 2014.
- Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. arXiv preprint:2111.02114, 2021.
- Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., Parikh, D., Gupta, S., and Taigman, Y. Make-a-video: Text-to-video generation without text-video data. arXiv preprint:2209.14792, 2022.
- Sinha, A., Song, J., Meng, C., and Ermon, S. D2c: Diffusion-decoding models for few-shot conditional generation. In Conference on Neural Information Processing Systems, 2021.
- Song, J., Meng, C., and Ermon, S. Denoising diffusion implicit models. arXiv preprint:2010.02502, 2020.
- Song, Y., Sohl-Dickstein, J., Kingma, D., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations, 2021.
- Tan, X., Chen, J., Liu, H., Cong, J., Zhang, C., Liu, Y., Wang, X., Leng, Y., Yi, Y., He, L., et al. NaturalSpeech: End-to-end text to speech synthesis with human-level quality. arXiv preprint:2205.04421, 2022.
- Vahdat, A., Kreis, K., and Kautz, J. Lsgm: Score-based generative modeling in latent space. In Conference on Neural Information Processing Systems, 2021.
- Wang, H. and Wang, D. Towards robust speech super-resolution. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:2058–2066, 2021.
- Wu, Y., Chen, K., Zhang, T., Hui, Y., Berg-Kirkpatrick, T., and Dubnov, S. Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation. arXiv preprint:2211:06687, 2022.
- Yamagishi, J., Veaux, C., MacDonald, K., et al. CSTR VCTK corpus: English multi-speaker corpus for cstr voice cloning toolkit. 2019.
- Yang, D., Yu, J., Wang, H., Wang, W., Weng, C., Zou, Y., and Yu, D. Diffsound: Discrete diffusion model for text-to-sound generation. arXiv preprint:2207.09983, 2022.
- Żelaszczyk, M. and Mańdziuk, J. Audio-to-image cross-modal generation. In International Joint Conference on Neural Networks, pp. 1–8. IEEE, 2022.
附录
A. 对比语言—音频预训练
作者沿用 Wu 等人提出的对比语言—音频预训练(CLAP)流程,学习文本与音频的相似度,并将二者投影到联合 latent space。训练集包括当时最大的公开数据集 LAION-Audio-K、以 T 模型把关键词扩展成 caption 的 AudioSet、AudioCaps 和 Clotho。LAION-Audio-K 含 个语言—音频对、共 小时;AudioSet 含 对、共 小时;AudioCaps 含 对、共 小时;Clotho 含 对、共 小时。这些数据覆盖多种自然声音、声效、音乐和人类活动。
给定音频样本 与文本数据 ,作者分别使用音频编码器和文本编码器提取 与 ,并令 。音频编码器基于 HTSAT,文本编码器基于 RoBERTa。用于训练这两个对比编码器的对称 cross-entropy loss 为:
其中, 是可学习的 temperature 参数, 是 batch size。
B. Latent Diffusion Model
AudioLDM 的 LDM 采用 StableDiffusion 的 UNet backbone。按照式(5),UNet 同时以时间步 与 CLAP embedding 为条件:作者先把时间步映射为一维 embedding,再与 拼接。由于条件向量只有一维,模型不使用 StableDiffusion 的 cross-attention,而直接通过 feature-wise linear modulation(FiLM)层把条件信息融合到 UNet 卷积块的 feature map 中。UNet 含四个 encoder block、一个 middle block 和四个 decoder block。以基础通道数 表示时,encoder 的通道维度为 ,decoder 次序相反,middle block 的通道维度为 。最后三个 encoder block 与最先三个 decoder block 均加入 attention block;每个 attention block 由两层 multi-head self-attention 和位于中间的一层 fully connected layer 组成,head 数等于 attention embedding 维度除以 。AudioLDM-S 取 ,AudioLDM-L 取 。前向过程使用 步,线性 noise schedule 从 增至 ;采样使用 步 DDIM。Classifier-free guidance 在文中所指公式中采用 。
C. Variational Autoencoder
作者使用卷积 VAE,把音频 的 mel-spectrogram 压缩到较小的连续空间 ;其中 分别是时、频维长度, 是 latent encoding 的通道数, 是 latent space 的压缩级别(下采样比例)。Encoder 与 decoder 均由堆叠卷积模块组成,因此 VAE encoder 能大致保持 mel-spectrogram 与 latent space 的空间对应关系,如 Figure 7 所示。每个模块由带卷积层和残差连接的 ResNet block 构成。Encoding 会均匀拆成 与 两部分,形状均为 ,分别表示 VAE latent space 的均值和方差。Decoder 的输入是随机 encoding ,其中 。生成时,decoder 根据生成的 latent 表征重建 mel-spectrogram。
训练目标使用三种损失:mel-spectrogram 重建损失、对抗损失和 Gaussian constraint loss。重建损失计算输入 与重建 mel-spectrogram 之间的 mean absolute error;对抗损失用于增强重建质量。具体而言,作者采用 PatchGAN 作为 discriminator:它把输入图像分成小 patch,并输出 logits 矩阵,逐个判断 patch 为真或假。
PatchGAN discriminator 的训练目标是增大正确识别真实 patch 的 logit,同时减小把伪造 patch 错认成真实 patch 的 logit。作者还对 VAE latent space 施加 Gaussian constraint,使 VAE 更倾向于学习连续、有结构的 latent space,而不是无组织空间;这有助于捕获数据底层结构,使重建更稳定、更准确。
VAE 使用 Adam optimizer,learning rate 为 ,batch size 为 ;音频数据包括 AudioSet、AudioCaps、Freesound 与 BBC SFX。作者实验三种压缩级别 ,对应 latent 通道数分别为 。三种 VAE 都在单张 NVIDIA RTX 3090 上训练至少 M 步。为稳定训练,前 K 步不使用对抗损失;数据增强使用 mixup。
Table 8 报告不同 下的 VAE 重建性能。三种设置的指标都与“GT Mel + Vocoder”相当,说明 autoencoder 能可靠地编码、解码 mel-spectrogram。
| 设置 | PSNR ↑ | SSIM ↑ | FD ↓ | IS ↑ | KL ↓ |
|---|---|---|---|---|---|
| GT Mel + Vocoder | 25.41 | 0.86 | 8.76 | 10.71 | 0.23 |
| Compression | 25.38 | 0.86 | 9.02 | 10.67 | 0.23 |
| Compression | 25.14 | 0.84 | 9.68 | 10.50 | 0.25 |
| Compression | 24.87 | 0.82 | 9.90 | 9.84 | 0.29 |
D. Vocoder
本文采用广泛用于语音波形生成的 HiFi-GAN 作为 vocoder。它包含两组 discriminator——multi-period discriminator 与 multi-scale discriminator——用于增强感知质量。作者在 AudioSet 上训练 vocoder 来合成音频波形。对采样率 Hz 的输入,提取 band mel-spectrogram,并沿用 HiFi-GAN V 的默认设置:window、FFT 与 hop size 分别为 ,,。AdamW optimizer 的两个参数为 与 ;learning rate 从 开始,decay 为 。模型使用 batch size ,在 张 NVIDIA 3090 上训练。作者在开源实现中发布了该预训练 vocoder。
E. 实验细节
数据处理。 AudioSet 与 AudioCaps 的音频时长为 秒,而 Freesound 与 BBC SFX 的时长通常更长。为避免过度使用长音频中往往重复的声音,作者只保留 Freesound 与 BBC SFX 每段音频的前 秒,并切分为十秒片段,最终得到 个十秒训练样本。即使 AudioCaps、BBC SFX 等数据集提供文本 caption,LDM 训练也完全不使用,只使用音频。所有数据重采样为 kHz、单声道,并 padding 到 秒。
配置。 每个 LDM 默认使用压缩级别 。AudioLDM-S 与 AudioLDM-L 分别在单张 NVIDIA RTX 3090 上训练 M 步,batch size 分别为 与 ,learning rate 均为 。AudioLDM-L-Full 在一张 NVIDIA A100 上训练 M 步,batch size 为 ,learning rate 为 ;为在 AudioCaps 上取得更好性能,评估前又在 AudioCaps 上微调 M 步。作者指出,GPU 稀缺迫使其限制 batch size,这可能限制了 AudioLDM 性能。相比之下,DiffSound 用 张 NVIDIA V100,每张卡 batch size 为 ;AudioGen 用 张 A100,总 batch size 为 。
主观评估。 作者随机选择 个样本: 个来自 AudioCaps, 个来自 AudioSet,另有 个真实录音作为 spam case。因此每个模型需依据相应文本描述生成 个音频样本。所有模型输出集中到同一文件夹,用随机标识匿名化。Table 9 给出问卷示例,参与者需要为每个音频文件填写最后两列。最终所有人工评审在 spam case 上的平均得分都高于 ,因此其评估结果被视为可靠。
| 文件名 | 文本描述 | 整体印象(1–100) | 与文本描述的相关性(1–100) |
|---|---|---|---|
| random_name_108029.wav | A man talking followed by lights scrapping on a wooden surface | 80 | 90 |
| random_name_108436.wav | Bicycle Music Skateboard Vehicle | 70 | 80 |
| random_name_116883.wav | A power tool drilling as rock music plays | 90 | 95 |
| … | … | … | … |
F. 微调的影响
| 模型 | 文本数据 | 使用 CLAP | 已微调 | FD ↓ | IS ↑ | KL ↓ | FAD ↓ |
|---|---|---|---|---|---|---|---|
| AudioLDM-S-Full-RoBERTa | ✓ | ✗ | ✗ | 34.28 | 3.53 | 3.44 | 6.96 |
| AudioLDM-S-Full-RoBERTa | ✓ | ✗ | ✓ | 32.13 | 4.02 | 3.25 | 5.89 |
| AudioLDM-S-Full | ✗ | ✓ | ✗ | 24.13 | 6.68 | 2.36 | 4.94 |
| AudioLDM-S-Full | ✗ | ✓ | ✓ | 23.47 | 7.57 | 1.98 | 2.32 |
| AudioLDM-L-Full | ✗ | ✓ | ✗ | 23.51 | 7.11 | 2.19 | 4.19 |
| AudioLDM-L-Full | ✗ | ✓ | ✓ | 23.31 | 8.13 | 1.59 | 1.96 |
Table 10 比较在评估集上微调与不微调的结果。各项指标均有所改善,这是预期现象,因为 AudioCaps 训练集与评估集分布相似。不过,在有限评估分布上取得更高性能,并不一定代表整体性能更好:能生成更广泛音频分布的模型,即使泛化能力更强,也可能在该评估集上表现更差。未来音频生成研究可着重建立与人类感知更一致的评估协议。
G. 计算效率比较
如 Figure 8 左图所示,不使用 classifier-free guidance 时,AudioLDM-S 可在十秒内生成八段十秒音频。使用 classifier-free guidance 后,在 个 DDIM step 下,AudioLDM-Small 仍能生成八段十秒音频。右图显示,不同 batch size 下本文模型都快于 DiffSound:AudioLDM 约用 秒即可生成八段十秒音频,而 DiffSound 需要超过 秒。由于 AudioGen 尚未开源,本文没有与其比较速度。
H. 局限性
本研究仍有若干局限值得后续探索。其一,模型采样率仍然不足,尤其影响音乐生成;探索 kHz 或 kHz 等更高保真采样率,可能改善生成质量。其二,AudioLDM 的各个模块分别训练,可能导致模块之间不对齐。例如,VAE 学到的 latent space 未必最适合 latent diffusion model。未来可通过端到端微调等方法更好地对齐各模块。
该方法潜在的负面影响包括技术或已发布模型被滥用,例如生成伪造声效以提供误导信息。此外,未来应限制敏感文本内容,以避免创建有害音频内容。
I. 示例
I.1 音频风格迁移
作者用所提出的浅层反向过程(见式(10)),展示 AudioLDM-S 的三个零样本音频风格迁移示例。Figure 9 从 drum beats 迁移到 ambient music:从左到右依次是源鼓点,以及在不同起点 下由文本 prompt ambient music 引导生成的六个样本。较小 (图左侧)生成结果更接近鼓点;最后一个样本取 时,更贴近文本输入 ambient music。类似地,Figure 10 展示以源音频 trumpet 和 prompt children singing 得到的七个样本;Figure 11 展示以源音频 sheep vocalization 和 prompt narration, monologue 得到的五个样本。
I.2 音频超分辨率
Figure 12 展示 AudioLDM-S 的四个零样本音频超分辨率案例:(1)violin;(2)sneezing sound from a woman;(3)baby crying;(4)female speech。输入样本(左)的采样率为 kHz,生成样本(中)与真实样本(右)均为 kHz。可视化表明,模型保留了低频部分(低于 kHz)的真实观测,同时用预训练 AudioLDM-S 生成高频缺失部分(– kHz);生成的高频信息与低频观测一致。
I.3 音频补绘
Figure 13 展示 AudioLDM-S 的四个零样本音频补绘样本。每个音频长 秒。“unprocessed” 行把真实样本的 – 秒内容移除,作为补绘输入;“inpainting result” 行展示使用与真实样本相同文本 prompt 引导生成的结果;“ground truth” 行给出真实样本以作比较。
Figure 14 用一个样本展示由不同文本 prompt 引导的音频补绘。给定顶行所示观测音频,分别用四个 prompt 引导补绘:(1)ambient music;(2)a man is speaking with bird calls in the background;(3)a cat is meowing;(4)raining with wind blowing。每个生成样本都保留了已观测音频信号,而生成内容可以由文本输入控制。
I.4 声学环境控制
Figure 15 展示 AudioLDM 能用文本描述控制生成样本的声学环境。四个样本使用相同随机种子,只改变文本 prompt;共同文本是 “A man is speaking in”,具体环境分别为 “a small room”、“a huge room”、“a huge room without background noise” 和 “a studio”。这些样本说明 AudioLDM 能捕获声学环境的细粒度文本描述,并控制混响或背景噪声等相应效果。
I.5 音乐属性控制
Figure 16 展示通过文本输入控制音乐特征时的生成样本。第一个样本由 “Theme music with bass drum” 生成;随后分别加入 “flute”、“fast, flute” 或 “flute in the background” 来改变文本输入,相应变化可从生成的 mel-spectrogram 中看到。这些样本说明 AudioLDM 能向音乐加入新乐器、调整速度,并控制前景—背景关系。
I.6 音高控制
Figure 17 展示 AudioLDM 控制生成样本音高的能力。音高是声效、音乐与语音的重要特征。这里共同文本为 “Sine wave with pitch”,具体文本分别为 “low”、“medium” 和 “high”。三个生成样本清楚呈现了文本控制的音高变化。
I.7 材料控制
Figure 18 展示 AudioLDM 控制发声材料的能力。四个样本使用共同动作 “hit”,但指定不同材料组合,例如 wooden object 与 wooden environment,或 metal object 与 wooden environment。
I.8 时间顺序控制
Figure 19 展示 AudioLDM 控制组合音频信号之间时间顺序的能力。当文本描述包含多个声效时,AudioLDM 能生成这些音频信号,并使它们之间的时间顺序与文本输入一致。
I.9 文本到音频生成
Figure 20 展示 AudioLDM-S 的四个文本到音频生成结果,覆盖自然环境声效、人类语音、人类活动以及物体交互产生的声音。
I.10 新颖音频生成
Figure 21 展示 AudioLDM-S 生成的四个新颖音频样本。它们的文本描述很少见,例如 “A wolf is singing a beautiful song.”;作者用这些样本展示 AudioLDM 的泛化能力。
I.11 音乐生成
Figure 22 展示 AudioLDM-S 生成的四个音乐样本。这里使用 AudioSet 标签作为音乐生成的文本描述,并能指定生成样本的音乐流派,例如 Classical music。