M3-071M3: Deep Learning FoundationsTraining Stability & Mixed PrecisionEasy
Mastery:

Training Stability & Mixed Precision: 解释 FP16、BF16、FP32 的差异与取舍。

📐 Mathematical Definition
FP16: 1+5+10;BF16: 1+8+7;FP32: 1+8+23 (sign+exp+mantissa)\text{FP16}:\ 1+5+10;\qquad \text{BF16}:\ 1+8+7;\qquad \text{FP32}:\ 1+8+23\ (\text{sign}+\text{exp}+\text{mantissa})
⚡ Executive Summary
Core Concept: FP32 精度高但慢;FP16 精度高但范围窄(易溢出);BF16 范围与 FP32 同但精度低;大模型训练首选 BF16。

📌 Key Takeaways

  • •
    FP16 指数 5 位(最大 65504)→ 溢出风险高
  • •
    BF16 指数 8 位(与 FP32 同范围)→ 无溢出但有效位少
  • •
    大模型训练用 BF16,推理可用 FP16/INT8

📐 Mathematical Derivations

数学机理:浮点数由 符号 + 指数 + 尾数 组成;<strong>指数位数决定动态范围</strong>(能表示的最大/最小值),<strong>尾数位数决定精度</strong>(有效数字位数)。<strong>FP32</strong>:8 位指数(范围约 1e-38~3e38)、23 位尾数(约 7 位十进制有效数字)。<strong>FP16</strong>:5 位指数(范围约 6e-5~65504)、10 位尾数(约 3 位有效数字)——<strong>范围窄</strong>是致命问题:梯度常小于 6e-5(下溢为 0)、激活/指数运算易超过 65504(上溢为 inf)。<strong>BF16</strong>:8 位指数(范围与 FP32 完全相同)、7 位尾数(约 2~3 位有效数字)——<strong>范围与 FP32 一致,故不会溢出</strong>,代价是精度较低(相对误差约 0.8%,而 FP16 约 0.1%)。<strong>取舍逻辑</strong>:深度学习对<strong>范围</strong>的敏感度远高于<strong>精度</strong>(梯度跨越多个量级、需避免溢出),故 BF16 的'同范围、低精度'恰好匹配需求;而 FP16 需要 loss scaling 手工补偿范围问题。<strong>硬件</strong>:Ampere 及以后的 GPU 原生支持 BF16(张量核心),且 BF16 与 FP32 的转换几乎无损(仅截断尾数),故大模型训练默认 BF16。

🏭 Production Trade-offs

深度剖析与工程权衡:① <strong>为什么范围比精度重要</strong>——训练中梯度的动态范围可跨越 1e-8~1e3,FP16 的下界 6e-5 会让小梯度直接下溢为 0('梯度消失'的假象);BF16 的下界约 1e-38,完全覆盖。② <strong>loss scaling 的机制</strong>——FP16 训练把 loss 乘以 scale(如 2^16)使梯度整体上移、避免下溢,更新前再除回来;scale 需动态调整(溢出时减半、长时间不溢出则翻倍)。BF16 因范围足够而无需此机制。③ <strong>精度损失的影响</strong>——BF16 的 0.8% 相对误差在'累加'场景(如 softmax 的 log-sum-exp、attention 的累加)可能放大;故关键算子(softmax、LayerNorm、loss)应保持 FP32 计算(混合精度)。④ <strong>推理侧的差异</strong>——推理对精度更敏感(误差会直接影响输出),故 FP16 推理(精度更高)比 BF16 更常见;而训练因有大量样本平均、噪声可容忍。⑤ <strong>与量化的关系</strong>——BF16/FP16 是'16 位浮点',INT8/INT4 是'低位整数',后者范围与精度都受限但吞吐更高;量化是推理优化的下一站。⑥ <strong>面试要点</strong>——被问'为什么大模型用 BF16 不用 FP16',核心答案是'<strong>范围与 FP32 一致、避免溢出,无需 loss scaling</strong>';只说'BF16 更省显存'是错的(两者都是 2 字节)。
⚠️ Common Interview Pitfalls
  • ✕
    以为 BF16 比 FP16 精度更高(恰恰相反,BF16 尾数更少)
  • ✕
    在 BF16 下仍保留 loss scaling(无必要)
🎯 Interviewer Follow-ups
  • ?
    为什么 FP16 需要 loss scaling 而 BF16 不需要?
  • ?
    BF16 的精度损失在哪些场景会显现?
📚

Associated Knowledge Base Guides & Mindmaps

Explore the comprehensive technical article, exam cards, and global architecture tree.

← PreviousM3-070: Loss Functions & Objectives: 解释 ranking loss(pairwise / listwise)与它在检索排序中的应用。📋Back to BankNext →M3-072: Training Stability & Mixed Precision: 解释 loss scaling 的机制与动态调整。