M1-035M1: Mathematics & Statistics FundamentalsNumerical StabilityHard
Mastery:

Numerical Stability: 列举深度学习中常见的数值不稳定来源,以及各自的缓解手段。

📐 Mathematical Definition
缓解: log⁡-sum-exp, ϵ 平滑, clipping, BF16, LN/RMSNorm\text{缓解}:\ \log\text{-sum-exp},\ \epsilon\ \text{平滑},\ \text{clipping},\ \text{BF16},\ \text{LN/RMSNorm}
⚡ Executive Summary
Core Concept: 来源:指数溢出、除零、log(0)、梯度爆炸/下溢、FP16 范围、方差过小归一化。

📌 Key Takeaways

  • •
    归一化层分母加 ε
  • •
    概率裁剪到 [1e-12, 1-1e-12]
  • •
    监控 grad norm 与 loss scale
  • •
    数值问题常表现为 loss NaN,需逐层定位

📐 Mathematical Derivations

按算子分类更系统:① <strong>指数类</strong>(softmax、sigmoid、exp、log-sum-exp)——溢出/下溢,缓解是减 max、分段计算、log1p;② <strong>除法与归一化</strong>(LayerNorm、BatchNorm、attention 的 1/√d)——除零或方差过小,缓解是加 ε(且 FP16 下 ε 需放大到 1e-3 量级);③ <strong>对数类</strong>(交叉熵、KL、log 概率)——log(0)=−inf,缓解是 clip 概率到 [1e-12, 1−1e-12] 或用 log 域直接计算;④ <strong>梯度类</strong>——爆炸(连乘 >1)或下溢(连乘 <1 或 FP16 范围),缓解是梯度裁剪、BF16、loss scaling、残差连接;⑤ <strong>累加类</strong>(大数求和、注意力 logits)——浮点累加误差,缓解是 Kahan 求和、FP32 累加。

🏭 Production Trade-offs

系统化的排查与预防:① <strong>定位手段</strong>——用 <code>torch.autograd.set_detect_anomaly(True)</code> 捕获产生 NaN 的具体算子;逐层 hook 监控激活与梯度的 min/max/norm;先看是哪一步(step)开始 NaN,再二分定位层。② <strong>结构性的稳定化</strong>——残差连接保证梯度至少有恒等通路(+1),归一化层把激活分布拉回稳定尺度,二者共同改善数值条件,这也是深层网络能训练的前提。③ <strong>训练层面的防线</strong>——梯度裁剪(clip by norm,LLM 常用 1.0)、BF16 替代 FP16、降低学习率、loss scaling 的动态调整。④ <strong>监控指标</strong>——grad norm 分布、update/param 比值(健康范围约 10⁻³)、loss scale 值的变化趋势,这些指标往往在 loss 变 NaN 之前就已异常。
⚠️ Common Interview Pitfalls
  • ✕
    只在 loss 变 NaN 后才排查(应监控 grad norm 提前预警)
  • ✕
    FP16 下沿用 FP32 的 ε 值(导致除零)
🎯 Interviewer Follow-ups
  • ?
    如何定位是哪一层产生 NaN?
  • ?
    为什么残差与归一化能改善数值条件?
📚

Associated Knowledge Base Guides & Mindmaps

Explore the comprehensive technical article, exam cards, and global architecture tree.

← PreviousM1-034: Numerical Stability: 解释为什么用 log1p(exp(-|z|)) 计算 BCE,而不是直接 log(1+exp(-z))。📋Back to BankNext →M1-036: 估计理论 (MLE/MAP): 定义最大似然估计(MLE),并给出一个完整例子。