M1-035M1: Mathematics & Statistics FundamentalsNumerical StabilityHard
Mastery:
Numerical Stability: 列举深度学习中常见的数值不稳定来源,以及各自的缓解手段。
📐 Mathematical Definition
⚡ Executive Summary
Core Concept: 来源:指数溢出、除零、log(0)、梯度爆炸/下溢、FP16 范围、方差过小归一化。
📌 Key Takeaways
- •归一化层分母加 ε
- •概率裁剪到 [1e-12, 1-1e-12]
- •监控 grad norm 与 loss scale
- •数值问题常表现为 loss NaN,需逐层定位
📐 Mathematical Derivations
按算子分类更系统:① <strong>指数类</strong>(softmax、sigmoid、exp、log-sum-exp)——溢出/下溢,缓解是减 max、分段计算、log1p;② <strong>除法与归一化</strong>(LayerNorm、BatchNorm、attention 的 1/√d)——除零或方差过小,缓解是加 ε(且 FP16 下 ε 需放大到 1e-3 量级);③ <strong>对数类</strong>(交叉熵、KL、log 概率)——log(0)=−inf,缓解是 clip 概率到 [1e-12, 1−1e-12] 或用 log 域直接计算;④ <strong>梯度类</strong>——爆炸(连乘 >1)或下溢(连乘 <1 或 FP16 范围),缓解是梯度裁剪、BF16、loss scaling、残差连接;⑤ <strong>累加类</strong>(大数求和、注意力 logits)——浮点累加误差,缓解是 Kahan 求和、FP32 累加。
🏭 Production Trade-offs
系统化的排查与预防:① <strong>定位手段</strong>——用 <code>torch.autograd.set_detect_anomaly(True)</code> 捕获产生 NaN 的具体算子;逐层 hook 监控激活与梯度的 min/max/norm;先看是哪一步(step)开始 NaN,再二分定位层。② <strong>结构性的稳定化</strong>——残差连接保证梯度至少有恒等通路(+1),归一化层把激活分布拉回稳定尺度,二者共同改善数值条件,这也是深层网络能训练的前提。③ <strong>训练层面的防线</strong>——梯度裁剪(clip by norm,LLM 常用 1.0)、BF16 替代 FP16、降低学习率、loss scaling 的动态调整。④ <strong>监控指标</strong>——grad norm 分布、update/param 比值(健康范围约 10⁻³)、loss scale 值的变化趋势,这些指标往往在 loss 变 NaN 之前就已异常。
⚠️ Common Interview Pitfalls
- ✕只在 loss 变 NaN 后才排查(应监控 grad norm 提前预警)
- ✕FP16 下沿用 FP32 的 ε 值(导致除零)
🎯 Interviewer Follow-ups
- ?如何定位是哪一层产生 NaN?
- ?为什么残差与归一化能改善数值条件?
📚
Associated Knowledge Base Guides & Mindmaps
Explore the comprehensive technical article, exam cards, and global architecture tree.