M6-047M6: Multimodal & Generative ModelsDiffusion Models Foundations (DDPM)Hard
Mastery:
Diffusion Models Foundations (DDPM): 解释扩散模型的推理加速路线。
📐 Mathematical Definition
⚡ Executive Summary
Core Concept: 减少步数(高阶求解器/蒸馏)、缓存复用、并行化、以及潜空间压缩(LDM)与量化。
📌 Key Takeaways
- •少步采样:DDIM/DPM-Solver(10~50 步)、一致性模型(1~4 步)
- •蒸馏:把多步教师蒸馏到少步学生(LCM/Turbo)
- •缓存:复用相邻步的特征(DeepCache)
- •潜空间:LDM 把计算搬到低维潜空间(省 10~100 倍)
📐 Mathematical Derivations
数学机理:<strong>五类加速路线</strong>。(1) <strong>少步求解器</strong>——把采样视为 ODE 求解,用<strong>高阶求解器</strong>(如 DPM-Solver、UniPC、Heun)在 10~50 步内达到 1000 步的质量。<strong>原理</strong>——ODE 的数值积分可用高阶方法(每步利用多阶导数信息);DDIM 是一阶,DPM-Solver 是二/三阶。<strong>收益</strong>——20~50 倍加速。(2) <strong>蒸馏(distillation)</strong>——把'多步教师'蒸馏到'少步学生':(a) <strong>渐进蒸馏</strong>(Progressive Distillation)——学生一步学教师两步;(b) <strong>一致性模型(Consistency Models)</strong>——训练模型使'同一 ODE 轨迹上的任意点都映射到同一结果',从而可 1~2 步生成;(c) <strong>LCM(Latent Consistency Model)</strong>、<strong>SDXL-Turbo</strong>(对抗蒸馏)——4 步甚至 1 步生成。<strong>收益</strong>——10~100 倍(但需额外训练,且可能损失多样性)。(3) <strong>缓存(caching)</strong>——相邻时间步的特征/中间结果高度相似,故可<strong>复用</strong>(如 DeepCache 复用 U-Net 的深层特征);<strong>收益</strong>——2~5 倍(几乎无损)。(4) <strong>潜空间(latent space)</strong>——<strong>Latent Diffusion(LDM)</strong> 把扩散过程搬到<strong>低维潜空间</strong>(如 64×64×4 而非 512×512×3),计算量降 10~100 倍;<strong>这是最大的加速</strong>(Stable Diffusion 的关键)。(5) <strong>量化与系统优化</strong>——(a) 权重量化(INT8/INT4);(b) 算子融合;(c) 批处理(多张图并行);(d) 编译(torch.compile)。<strong>收益</strong>——1.5~3 倍(无损)。<strong>其他</strong>——(a) <strong>并行采样</strong>(Parallel sampling,如 ParaDiGMS 用 Picard 迭代并行化);(b) <strong>减少分辨率</strong>(生成低分辨率再超分);(c) <strong>CFG 的优化</strong>(如 CFG 需两次前向,可用'CFG distillation'合并为一次)。<strong>优先级</strong>——(a) <strong>先做潜空间</strong>(LDM,收益最大且已成熟);(b) <strong>再用少步求解器</strong>(免费加速);(c) <strong>缓存</strong>(几乎无损);(d) <strong>蒸馏</strong>(收益大但需训练);(e) <strong>量化/系统</strong>(工程手段)。<strong>权衡</strong>——(a) <strong>少步求解器</strong>——无损(只是数值求解方式不同);(b) <strong>蒸馏</strong>——可能损失多样性(因为学生模型被'压缩');(c) <strong>缓存</strong>——几乎无损;(d) <strong>量化</strong>——可能影响细节质量。<strong>度量</strong>——(a) <strong>每张图的延迟/成本</strong>;(b) <strong>FID/CLIP-score</strong>(质量);(c) <strong>多样性</strong>(蒸馏后常下降)。
🏭 Production Trade-offs
深度剖析与工程权衡:① <strong>'潜空间是最大收益'</strong>——LDM 把计算从像素空间搬到低维潜空间,收益 10~100 倍且质量损失小;这是'扩散模型能实用'的关键。② <strong>'少步求解器是无损加速'</strong>——它不改变模型(只是更好的数值方法),故应默认使用。③ <strong>'蒸馏损失多样性'</strong>——一致性模型/对抗蒸馏可 1~4 步生成,但多样性常下降(因为'多步探索'被压缩);故需在'速度 vs 多样性'间权衡。④ <strong>'缓存几乎无损'</strong>——利用相邻步特征的相似性,是'免费'的加速;DeepCache 等即此。⑤ <strong>'CFG 的双倍成本'</strong>——CFG 需'条件 + 无条件'两次前向(成本 ×2);故有'CFG distillation'(把两次合并为一次)的加速手段。⑥ <strong>面试要点</strong>——被问'扩散怎么加速',应给出'<strong>五类路线(少步求解器 / 蒸馏 / 缓存 / 潜空间 / 量化系统)+ 优先级(先潜空间)+ 各路的无损性</strong>';能指出'CFG 需两次前向是常见成本'是深度理解的标志。
⚠️ Common Interview Pitfalls
- ✕只用更多步数提升质量(应优化求解器)
- ✕忽略 CFG 的双倍前向成本
🎯 Interviewer Follow-ups
- ?一致性模型为什么能 1 步生成?
- ?哪条路线收益最大?
📚
Associated Knowledge Base Guides & Mindmaps
Explore the comprehensive technical article, exam cards, and global architecture tree.