Audio-Zero
Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning
完整覆盖 16 页、4 幅原论文图、Table 1–5、公式(1)–(11)、Impact Statement、附录 A–E 与完整参考文献,重点呈现 Listening–Attribution 双阶段 self-play、可验证游戏奖励和交替 GRPO。
- Audio Reasoning
- Self-Play
- GRPO
这里只收录你明确指定,且已完成逐页校对与覆盖验证的中文全文译稿。普通论文仍保留在资料摘要与阅读进度中。
Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning
完整覆盖 16 页、4 幅原论文图、Table 1–5、公式(1)–(11)、Impact Statement、附录 A–E 与完整参考文献,重点呈现 Listening–Attribution 双阶段 self-play、可验证游戏奖励和交替 GRPO。
ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
完整覆盖 8 页、167 个来源块、Figure 1、Table 1–6、11 个编号公式、作者脚注与 55 条参考文献,重点呈现 Zipformer 文本/向量场双骨干、Average Upsampling、masked CFM 与 4 NFE flow distillation。
AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation
完整覆盖 10 页、184 个来源块、Figure 1–3、Table 1–7、公式(1)–(11)、致谢与 78 条参考文献,重点呈现 audio-video 互补时间掩码、文本 cross-attention、CFM + 中间层 CTC 以及 multimodal CFG。
Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks
完整覆盖 8 页、116 个来源块、Figure 1–4、Table 1、编号公式(1)–(16)、致谢与 17 条参考文献,重点呈现 blank、路径折叠映射、forward-backward 边缘化及 maximum-likelihood 训练。
Training language models to follow instructions with human feedback
完整覆盖 68 页、420 个来源块、Figure 1–50、Table 1–14、2 个编号公式、全部附录与 91 条参考文献,重点呈现 SFT、reward model 与 PPO 三阶段 RLHF 流程及其对 helpfulness、truthfulness、toxicity 的影响。
LLaMA: Open and Efficient Foundation Language Models
完整覆盖 27 页、323 个来源块、Figure 1–3、Table 1–16、5 个 Algorithm、全部附录与 82 条参考文献,重点呈现 compute-optimal scaling、公开数据配方、训练基础设施与模型规模之间的取舍。
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
完整覆盖 52 页、294 个来源块、Figure 1–5、Table 1–37、47 个编号公式、全部附录与 62 条参考文献,重点呈现 MLA 的 KV cache 压缩、DeepSeekMoE 的细粒度专家与 shared expert,以及训练和推理效率。
GIVT: Generative Infinite-Vocabulary Transformers
完整覆盖 32 页、219 个来源块、Figure 1–21、Table 1–8、2 个编号公式、Algorithm、全部附录与 79 条参考文献,重点呈现无限词表连续 token、Gaussian mixture 输出分布及 causal/MaskGIT 两种生成方式。
Autoregressive Image Generation without Vector Quantization
完整覆盖 16 页、174 个来源块、Figure 1–8、Table 1–6、3 个编号公式、Algorithm、全部附录与 56 条参考文献,重点呈现用 Diffusion Loss 建模连续 token 条件分布、无需 vector quantization 的 autoregressive image generation。
Scalable Diffusion Models with Transformers
完整覆盖 25 页、174 个来源块、Figure 1–33、Table 1–6、全部数学表达、Appendix A–D 与 63 条参考文献,重点呈现 DiT block、adaLN-Zero 与计算量驱动的 scaling 规律。
AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
完整覆盖 25 页、228 个来源块、Figure 1–22、Table 1–10、14 个编号公式、Appendix A–I 与 63 条参考文献,重点呈现 CLAP 条件模态替换、连续 VAE latent 与统一零样本音频操作。
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
完整覆盖 27 页、186 个来源块、Figure 1–5、10 张编号表与 1 张志愿者表、22 个编号公式、Algorithm、Appendix A–D 与 51 条参考文献,重点呈现从 KL 约束 RLHF 到直接偏好分类损失的闭式重参数化。
Luna-TTS Family Technical Report
完整覆盖 21 页、283 个源块、Figure 1–2、Table 1–15、3 个编号公式与 7 个无编号展示公式、作者说明及 86 条参考文献,重点呈现 RVQ 网格 masked diffusion、fully parallel 与 block-causal streaming 两种工作点,以及 trajectory-aware GRPO。
DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation
完整覆盖 21 页、Figure 1–9、Table 1–8、公式(1)–(31)、Algorithm 1、3 个脚注、Appendix A–C 与 62 条参考文献,重点呈现 patch-based continuous-token autoregression、LocDiT outpainting、LM guidance 与 ODE-compatible temperature sampling。
dots.tts Technical Report
完整覆盖 22 页、Figure 1–2、Table 1–5、15 个编号公式与 13 个无编号展示公式、作者贡献及 45 条实际引用文献,重点呈现 AudioVAE、full-history AR-Flow、Self-corrective alignment 与 CFG-aware MeanFlow。
Qwen-Audio-VAE Technical Report
完整覆盖 15 页、Figure 1–4、Table 1–14、公式(1)–(2)、Core Contributors、Appendix 7.1–7.4 与 28 条参考文献,重点呈现 12.5 Hz continuous latent、500 万小时多领域训练与 3.62× 编码加速。
Stable Audio 3
完整覆盖 26 页、Figure 1–14、Table 1–14、公式(1)–(12)、2 个脚注与 96 条参考文献,重点呈现 SAME、原生变长训练、Adversarial Post-Training 与 8-step Ping-Pong sampling。
Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens
完整覆盖 45 页、66 个编号公式与 74 个无编号展示公式、Figure 1–8、Table 1–5、4 个脚注及 128 条参考文献,重点呈现 Locodec 的 token 空间塑形与 MP-ELD 的多路径 residual CFG。
ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling
完整覆盖 20 页正文、Figure 1–5、Table 1–11、公式(1)–(6)、Algorithm 1–2、Appendix A–G 与 45 条参考文献,重点呈现渐进数据构建、三阶段训练与渐进引导采样。
An Introduction to Flow Matching and Diffusion Models
84 页中文全文译稿,完整覆盖 Flow Matching、Diffusion Models、Score Matching、classifier-free guidance、Diffusion Transformer、VAE、离散 CTMC、附录与参考文献。
Qwen-Audio-3.0-Gen-Preview Technical Report
完整覆盖正文、原始系统图、Table 1–11、公式(1)–(6)、脚注、致谢与参考文献,重点呈现统一混合波形生成、rich-timeline 条件与共享连续 VAE。
Qwen3-TTS Technical Report
完整覆盖 14 页正文、双 speech tokenizer、dual-track Language Model、Figure 1–3、Table 1–10、作者脚注与 43 条参考文献,并保留报告中的两处原文数据不一致。
Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
完整覆盖正文、公式、原论文图表、Algorithm、参考文献与 Supplementary A–E,重点呈现 Rectified Flow 的 timestep sampling 与 MM-DiT 架构。