M3-073M3: Deep Learning FoundationsTraining Stability & Mixed PrecisionMedium
Mastery:
Training Stability & Mixed Precision: 训练出现 loss NaN 时如何排查?
📐 Mathematical Definition
⚡ Executive Summary
Core Concept: 按'先定位首次出现的位置,再分算子/数据/超参'排查:softmax 溢出、除零、lr 过大、脏数据、FP16 上溢。
📌 Key Takeaways
- •先定位首个 NaN 的 step 与算子(hook 前向/反向)
- •常见源:softmax 未减 max、LN 方差为 0、lr 过大、FP16 溢出
- •用 BF16、加 ε、降 lr、裁剪、跳过坏批次
📐 Mathematical Derivations
数学机理:NaN 的产生源头可穷举:<strong>0/0</strong>(如归一化时分母为 0)、<strong>∞/∞</strong>、<strong>log(0)=−∞</strong>(交叉熵在预测为 0 时)、<strong>exp(大数)=∞</strong>(softmax/激活未减 max)、<strong>sqrt(负数)</strong>(数值误差导致方差略负)。一旦某处产生 NaN,反向传播中的乘法/加法会让 NaN <strong>传染</strong>到所有相关梯度(NaN·x=NaN),进而污染全部参数——故'看到一个 NaN'不等于'源头在那里'。<strong>排查策略</strong>是'<strong>先定位首次出现</strong>':(1) 记录每步的 loss,找到第一个 NaN 的 step;(2) 用 <code>torch.autograd.set_detect_anomaly(True)</code> 让 PyTorch 在反向遇到 NaN 时抛出异常并打印产生 NaN 的具体算子(代价是速度变慢,仅用于调试);(3) 用 hook 记录每层前向输出与反向梯度的 min/max/是否含 NaN,二分定位到具体层。<strong>定位后按类别修复</strong>:(a) <strong>算子层</strong>——softmax 加 max-subtraction、LN/BN 的 ε 调大、log 前加 eps、用 stable 的 CE;(b) <strong>超参层</strong>——降 lr、加长 warmup、加梯度裁剪、增大 Adam 的 ε;(c) <strong>数据层</strong>——检查异常样本(超长序列、全 padding、重复 token);(d) <strong>精度层</strong>——从 FP16 换 BF16、关键算子保 FP32。
🏭 Production Trade-offs
深度剖析与工程权衡:① <strong>NaN vs inf 的区别</strong>——inf 通常是上溢(大数相乘/exp),NaN 通常是 0/0 或 inf−inf;先区分类型可缩小范围。② <strong>前向 NaN vs 反向 NaN</strong>——前向 NaN 说明激活/损失异常(算子或数据问题);前向正常但反向 NaN 说明梯度计算异常(如除以极小的方差、梯度爆炸)。分别用前向/反向 hook 定位。③ <strong>常见'隐蔽'源</strong>——(a) attention 的 mask 用 −inf 时若整行被 mask 则 softmax 得 0/0(需用 −1e9 或加保护);(b) 稀疏数据的 LN 方差为 0;(c) 混合精度下 loss scale 未正确 unscale。④ <strong>恢复策略</strong>——定位并修复后,若已污染参数,需回滚到 NaN 之前的 checkpoint 重训(用坏梯度更新过的参数无法自动恢复)。⑤ <strong>预防性设计</strong>——使用 BF16、稳定算子、梯度裁剪、skip-nan 机制(检测到 NaN 时跳过该步);这些是现代训练框架的标配。⑥ <strong>面试要点</strong>——被问'loss 变 NaN 怎么办',应给出'<strong>先定位首个 NaN + 区分前向/反向 + 按算子/超参/数据/精度四类排查</strong>'的结构化流程,并提到 set_detect_anomaly 这一具体工具;回答'调小 lr'就结束是明显不足。
⚠️ Common Interview Pitfalls
- ✕看到 NaN 就直接调 lr(未定位源头,可能反复出现)
- ✕忽略 mask 全为 −inf 导致的 0/0
🎯 Interviewer Follow-ups
- ?如何用 torch.autograd.set_detect_anomaly 定位?
- ?为什么 NaN 会'传染'整个 batch?
📚
Associated Knowledge Base Guides & Mindmaps
Explore the comprehensive technical article, exam cards, and global architecture tree.