M1-031M1: Mathematics & Statistics FundamentalsNumerical StabilityEasy
Mastery:

Numerical Stability: 为什么 softmax 必须先减最大值?

📐 Mathematical Definition
softmax(x)i=exi−max⁡jxj∑kexk−max⁡jxj\mathrm{softmax}(x)_i=\frac{e^{x_i-\max_j x_j}}{\sum_k e^{x_k-\max_j x_j}}
⚡ Executive Summary
Core Concept: exp 对大正数溢出为 inf;减去最大值后指数最大为 0,结果不变但数值安全。

📌 Key Takeaways

  • •
    数学上等价,数值上必需
  • •
    同理 log-sum-exp 用于交叉熵

📐 Mathematical Derivations

数学等价性来自 softmax 的<strong>平移不变性</strong>:softmax(x+c)=softmax(x),因为分子分母同乘 e^c 后约掉。因此减 max 不改变输出,只改变中间计算的数值范围。未减 max 时,若 x 中有元素 ≥ 89(FP32 下 exp(88.7)≈3.4×10³⁸ 接近 float32 上限 3.4×10³⁸),e^{xᵢ} 直接溢出为 inf,随后 inf/inf=NaN,梯度传播中断;在 FP16 下阈值更低(exp(11.1)≈65504 即 FP16 上限),所以混合精度训练中这个问题更频繁。减去 max 后,最大的指数为 e⁰=1,其余 ≤1,求和结果落在 [1, n] 区间,完全安全。

🏭 Production Trade-offs

两个进阶要点:① <strong>减 max 不是唯一选择</strong>——减任何常数都数学等价,减 max 只是保证不溢出;某些实现减 max 是为了配合'在线 softmax'(FlashAttention)的分块计算,此时维护一个<strong>运行最大值</strong> m 并随时修正历史累加和 l。② <strong>log 域的配套</strong>——计算 log(softmax(x)) 时应直接用 <code>x - logsumexp(x)</code>,而非 <code>log(softmax(x))</code>(后者在概率极小时 log(0)=−inf)。框架提供的 <code>log_softmax</code>、<code>cross_entropy</code>、<code>bce_with_logits</code> 都内置了这些技巧,这也是为什么<strong>绝不应该</strong>先 sigmoid/softmax 再取 log。③ FP16 下还需注意:减 max 后的和可能仍很小(如全部为 −100 时 e^{−100}≈3.7×10⁻⁴⁴ 在 FP16 下下溢为 0),此时应全程在 FP32 累加。
⚠️ Common Interview Pitfalls
  • ✕
    先 softmax 再取 log(数值灾难)
  • ✕
    在 FP16 下用 FP16 累加 softmax 分母
🎯 Interviewer Follow-ups
  • ?
    log-sum-exp 技巧的写法?
  • ?
    FP16 下还有哪些常见溢出点?
📚

Associated Knowledge Base Guides & Mindmaps

Explore the comprehensive technical article, exam cards, and global architecture tree.

← PreviousM1-030: Convex Optimization & KKT: 解释次梯度与近端算子(proximal operator),以及它们在 L1 优化中的作用。📋Back to BankNext →M1-032: Numerical Stability: 写出 log-sum-exp 技巧,并说明它解决了什么问题。