跳转到内容

输入关键词开始搜索

    论文译库

    这里只收录你明确指定,且已完成逐页校对与覆盖验证的中文全文译稿。普通论文仍保留在资料摘要与阅读进度中。

    23完成译稿

    Audio-Zero

    Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning

    完整覆盖 16 页、4 幅原论文图、Table 1–5、公式(1)–(11)、Impact Statement、附录 A–E 与完整参考文献,重点呈现 Listening–Attribution 双阶段 self-play、可验证游戏奖励和交替 GRPO。

    • Audio Reasoning
    • Self-Play
    • GRPO

    ZipVoice

    ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

    完整覆盖 8 页、167 个来源块、Figure 1、Table 1–6、11 个编号公式、作者脚注与 55 条参考文献,重点呈现 Zipformer 文本/向量场双骨干、Average Upsampling、masked CFM 与 4 NFE flow distillation。

    • Text-to-Speech
    • Flow Matching
    • Zipformer

    AlignDiT

    AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation

    完整覆盖 10 页、184 个来源块、Figure 1–3、Table 1–7、公式(1)–(11)、致谢与 78 条参考文献,重点呈现 audio-video 互补时间掩码、文本 cross-attention、CFM + 中间层 CTC 以及 multimodal CFG。

    • Multimodal Speech
    • Diffusion Transformer
    • Audio-Visual Alignment

    CTC

    Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks

    完整覆盖 8 页、116 个来源块、Figure 1–4、Table 1、编号公式(1)–(16)、致谢与 17 条参考文献,重点呈现 blank、路径折叠映射、forward-backward 边缘化及 maximum-likelihood 训练。

    • Sequence Labelling
    • Speech Recognition
    • Dynamic Programming

    InstructGPT

    Training language models to follow instructions with human feedback

    完整覆盖 68 页、420 个来源块、Figure 1–50、Table 1–14、2 个编号公式、全部附录与 91 条参考文献,重点呈现 SFT、reward model 与 PPO 三阶段 RLHF 流程及其对 helpfulness、truthfulness、toxicity 的影响。

    • RLHF
    • Instruction Following
    • Alignment

    LLaMA

    LLaMA: Open and Efficient Foundation Language Models

    完整覆盖 27 页、323 个来源块、Figure 1–3、Table 1–16、5 个 Algorithm、全部附录与 82 条参考文献,重点呈现 compute-optimal scaling、公开数据配方、训练基础设施与模型规模之间的取舍。

    • Foundation Model
    • Scaling Laws
    • Pretraining

    DeepSeek-V2

    DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

    完整覆盖 52 页、294 个来源块、Figure 1–5、Table 1–37、47 个编号公式、全部附录与 62 条参考文献,重点呈现 MLA 的 KV cache 压缩、DeepSeekMoE 的细粒度专家与 shared expert,以及训练和推理效率。

    • Mixture of Experts
    • MLA
    • Efficient Inference

    GIVT

    GIVT: Generative Infinite-Vocabulary Transformers

    完整覆盖 32 页、219 个来源块、Figure 1–21、Table 1–8、2 个编号公式、Algorithm、全部附录与 79 条参考文献,重点呈现无限词表连续 token、Gaussian mixture 输出分布及 causal/MaskGIT 两种生成方式。

    • Continuous Tokens
    • Image Generation
    • Gaussian Mixture

    MAR

    Autoregressive Image Generation without Vector Quantization

    完整覆盖 16 页、174 个来源块、Figure 1–8、Table 1–6、3 个编号公式、Algorithm、全部附录与 56 条参考文献,重点呈现用 Diffusion Loss 建模连续 token 条件分布、无需 vector quantization 的 autoregressive image generation。

    • Autoregressive Generation
    • Diffusion Loss
    • Continuous Tokens

    DiT

    Scalable Diffusion Models with Transformers

    完整覆盖 25 页、174 个来源块、Figure 1–33、Table 1–6、全部数学表达、Appendix A–D 与 63 条参考文献,重点呈现 DiT block、adaLN-Zero 与计算量驱动的 scaling 规律。

    • Diffusion Transformer
    • Scaling
    • Image Generation

    AudioLDM

    AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

    完整覆盖 25 页、228 个来源块、Figure 1–22、Table 1–10、14 个编号公式、Appendix A–I 与 63 条参考文献,重点呈现 CLAP 条件模态替换、连续 VAE latent 与统一零样本音频操作。

    • Text-to-Audio
    • Latent Diffusion
    • CLAP

    DPO

    Direct Preference Optimization: Your Language Model is Secretly a Reward Model

    完整覆盖 27 页、186 个来源块、Figure 1–5、10 张编号表与 1 张志愿者表、22 个编号公式、Algorithm、Appendix A–D 与 51 条参考文献,重点呈现从 KL 约束 RLHF 到直接偏好分类损失的闭式重参数化。

    • Preference Optimization
    • RLHF
    • Alignment

    Luna-TTS

    Luna-TTS Family Technical Report

    完整覆盖 21 页、283 个源块、Figure 1–2、Table 1–15、3 个编号公式与 7 个无编号展示公式、作者说明及 86 条参考文献,重点呈现 RVQ 网格 masked diffusion、fully parallel 与 block-causal streaming 两种工作点,以及 trajectory-aware GRPO。

    • Text-to-Speech
    • Masked Diffusion
    • Streaming TTS

    DiTAR

    DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation

    完整覆盖 21 页、Figure 1–9、Table 1–8、公式(1)–(31)、Algorithm 1、3 个脚注、Appendix A–C 与 62 条参考文献,重点呈现 patch-based continuous-token autoregression、LocDiT outpainting、LM guidance 与 ODE-compatible temperature sampling。

    • Speech Generation
    • Continuous Tokens
    • Diffusion Transformer

    dots.tts

    dots.tts Technical Report

    完整覆盖 22 页、Figure 1–2、Table 1–5、15 个编号公式与 13 个无编号展示公式、作者贡献及 45 条实际引用文献,重点呈现 AudioVAE、full-history AR-Flow、Self-corrective alignment 与 CFG-aware MeanFlow。

    • Text-to-Speech
    • Continuous AR
    • Flow Matching

    Qwen-Audio-VAE

    Qwen-Audio-VAE Technical Report

    完整覆盖 15 页、Figure 1–4、Table 1–14、公式(1)–(2)、Core Contributors、Appendix 7.1–7.4 与 28 条参考文献,重点呈现 12.5 Hz continuous latent、500 万小时多领域训练与 3.62× 编码加速。

    • Audio VAE
    • Continuous Latent
    • Fast Encoding

    Stable Audio 3

    Stable Audio 3

    完整覆盖 26 页、Figure 1–14、Table 1–14、公式(1)–(12)、2 个脚注与 96 条参考文献,重点呈现 SAME、原生变长训练、Adversarial Post-Training 与 8-step Ping-Pong sampling。

    • Audio Generation
    • Variable Length
    • Adversarial Post-Training

    Locodec / MP-ELD

    Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens

    完整覆盖 45 页、66 个编号公式与 74 个无编号展示公式、Figure 1–8、Table 1–5、4 个脚注及 128 条参考文献,重点呈现 Locodec 的 token 空间塑形与 MP-ELD 的多路径 residual CFG。

    • Speech Generation
    • Continuous Tokens
    • Flow Matching

    ControlAudio

    ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling

    完整覆盖 20 页正文、Figure 1–5、Table 1–11、公式(1)–(6)、Algorithm 1–2、Appendix A–G 与 45 条参考文献,重点呈现渐进数据构建、三阶段训练与渐进引导采样。

    • Text-to-Audio
    • Timing Control
    • Progressive Diffusion

    MIT 6.S184 · Flow Matching

    An Introduction to Flow Matching and Diffusion Models

    84 页中文全文译稿,完整覆盖 Flow Matching、Diffusion Models、Score Matching、classifier-free guidance、Diffusion Transformer、VAE、离散 CTMC、附录与参考文献。

    • Flow Matching
    • Diffusion Models
    • CTMC

    Qwen-Audio 3.0 Gen Preview

    Qwen-Audio-3.0-Gen-Preview Technical Report

    完整覆盖正文、原始系统图、Table 1–11、公式(1)–(6)、脚注、致谢与参考文献,重点呈现统一混合波形生成、rich-timeline 条件与共享连续 VAE。

    • Audio DiT
    • rich-timeline
    • Unified Audio Generation

    Qwen3-TTS

    Qwen3-TTS Technical Report

    完整覆盖 14 页正文、双 speech tokenizer、dual-track Language Model、Figure 1–3、Table 1–10、作者脚注与 43 条参考文献,并保留报告中的两处原文数据不一致。

    • Text-to-Speech
    • Voice Cloning
    • Streaming TTS

    SD3 / MM-DiT

    Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

    完整覆盖正文、公式、原论文图表、Algorithm、参考文献与 Supplementary A–E,重点呈现 Rectified Flow 的 timestep sampling 与 MM-DiT 架构。

    • Rectified Flow
    • MM-DiT
    • Image Synthesis