M1-073M1: Mathematics & Statistics FundamentalsNumerical StabilityHard
Mastery:
Numerical Stability: 解释 FP32 主权重(master weights)与它在混合精度训练中的作用。
📐 Mathematical Definition
⚡ Executive Summary
Core Concept: 保留一份 FP32 参数副本用于更新,前向反向用低精度;防止小更新被舍入吞掉。
📌 Key Takeaways
- •更新量常远小于参数本身(约 10⁻⁴ 倍)
- •FP16 下 θ+Δ 可能等于 θ(舍入吞掉更新)
📐 Mathematical Derivations
问题的数学本质:参数更新的相对幅度通常很小——经验上 ‖Δθ‖/‖θ‖≈10⁻³ 到 10⁻⁴。而 FP16 只有 10 位尾数(约 3–4 位十进制有效数字),当一个数与其增量之比超过 2¹¹ 时,加法结果会被舍入回原值(<strong>增量被完全吞掉</strong>)。例如 θ=1.0(FP16 下精度约 2⁻¹⁰≈0.001),若 Δθ=10⁻⁵,则 θ+Δθ 在 FP16 下仍为 1.0——参数<strong>永不更新</strong>。<strong>FP32 主权重的解法</strong>:始终保留一份 FP32 参数副本 θ_FP32,优化器在 FP32 中完成 <code>θ_FP32 ← θ_FP32 − η·ĝ</code>(FP32 有 23 位尾数,可表示 10⁻⁷ 级增量),每次前向传播前再把 θ_FP32 转换为 FP16/BF16 供计算使用。这样既享受低精度的速度与显存优势,又保证更新的数值精度。
🏭 Production Trade-offs
实践要点:① <strong>显存代价</strong>——FP32 主权重额外占用 4 字节/参数(相对 FP16 的 2 字节,即多 50% 的参数存储);这是混合精度训练'省显存'效果不如理论值(2×)的原因之一。② <strong>BF16 也需要</strong>——尽管 BF16 动态范围与 FP32 相同(不会下溢),但其尾数只有 7 位(精度更低),故同样需要 FP32 主权重来保证更新精度。③ <strong>与 loss scaling 的分工</strong>——两者解决<strong>不同</strong>问题:loss scaling 防止<strong>梯度下溢</strong>(FP16 下小梯度变 0),FP32 主权重防止<strong>参数更新被舍入吞掉</strong>(θ+Δθ=θ);BF16 通常不需要 loss scaling(范围大)但仍需主权重(精度低)。④ <strong>实现</strong>——PyTorch AMP 的 <code>autocast</code> + <code>GradScaler</code> 自动处理;FSDP/ZeRO 的优化器状态分片也是基于 FP32 主权重做切分。⑤ <strong>全 BF16 训练</strong>——近年有工作尝试完全不用主权重(直接 BF16 更新),依赖更大的 batch 与更多步数补偿精度损失,但主流仍保留 FP32 主权重。
⚠️ Common Interview Pitfalls
- ✕认为 BF16 不需要 FP32 主权重(精度仍不足)
- ✕把 loss scaling 与主权重混为一谈(解决不同问题)
🎯 Interviewer Follow-ups
- ?为什么 BF16 也建议保留 FP32 主权重?
- ?与 loss scaling 的分工是什么?
📚
Associated Knowledge Base Guides & Mindmaps
Explore the comprehensive technical article, exam cards, and global architecture tree.