M3-060M3: Deep Learning FoundationsGradient Vanishing & ExplosionHard
Mastery:

Gradient Vanishing & Explosion: 解释 loss spike 现象与梯度范数的关系及缓解。

📐 Mathematical Definition
spike: Lt≫median(L);precursor: ∥gt∥≫median(∥g∥)\text{spike}:\ \mathcal{L}_t\gg\text{median}(\mathcal{L});\qquad \text{precursor}:\ \left\|g_t\right\|\gg\text{median}(\left\|g\right\|)
⚡ Executive Summary
Core Concept: 训练中 loss 突然上升(常伴 grad_norm 尖峰),多由特定数据批次或数值问题引起;用回滚、跳过批次、降 lr、裁剪缓解。

📌 Key Takeaways

  • •
    grad_norm 尖峰通常先于 loss 突增(预警窗口)
  • •
    常见诱因:脏数据、异常长序列、注意力数值溢出
  • •
    缓解:回滚到前一 checkpoint、跳过该批次、降 lr

📐 Mathematical Derivations

数学机理:<strong>loss spike</strong> 指训练中 loss 突然上升一到数个量级(如从 2.0 跳到 50),随后可能恢复或持续恶化。其因果链通常是:某个数据批次(含异常长序列、重复/乱码文本、极端分布样本)产生异常大的梯度 → 参数被大幅推离当前极小值区域 → loss 升高。<strong>grad_norm 的预警作用</strong>来自时序:梯度先异常(第 t 步)→ 参数更新(第 t 步末)→ loss 升高(第 t+1 步),故 grad_norm 尖峰<strong>早于</strong> loss 突增一步,形成预警窗口。<strong>为什么能自行恢复</strong>:优化器(尤其带动量)会继续沿原方向把参数拉回,且后续正常批次的梯度会把参数推回盆地;但若参数被推到'坏区域'(如进入 ReLU 全负区、LN 统计失真),则可能无法恢复、需回滚。<strong>数值性诱因</strong>:FP16 下注意力 softmax 的 exp 溢出、LayerNorm 的除零、除以极小值等,都会产生 inf/NaN 或极大梯度。

🏭 Production Trade-offs

深度剖析与工程权衡:① <strong>标准处置流程</strong>——(a) 保存'回滚点'(如每 N 步的 checkpoint);(b) 检测到 spike 时回滚到 spike 前;(c) 跳过引起 spike 的数据批次(或该 batch 内的问题序列);(d) 降低 lr 或提高 grad clip 的严格度;(e) 若频繁发生,检查数据清洗与数值稳定性(用 BF16、稳定 softmax)。② <strong>数据侧根治</strong>——统计'每批的最大序列长度/重复率/困惑度',剔除异常;很多 loss spike 源于少数超长序列或数据管道 bug。③ <strong>优化器侧的缓解</strong>——增大 Adam 的 ε(吸收小梯度噪声)、降低 β₂(让 v 更快响应)、缩短 lr warmup 后的峰值 lr;也可用'loss 异常时跳过该步'的工程手段(如某些框架的 skip-nan 机制)。④ <strong>与 μP 的关系</strong>——μP 通过让更新量级与宽度解耦,显著降低大规模训练的 spike 频率;这是 μP 的实践卖点之一。⑤ <strong>z-loss 等技巧</strong>——在 softmax 上加 z-loss(惩罚 log-sum-exp 偏离 0)可稳定 logits、减少 spike;PaLM 等使用了这一技巧。⑥ <strong>面试要点</strong>——若被问'训练 loss 突然爆炸怎么办',应给出<strong>分步处置 + 根因分类(数据 vs 数值 vs 优化器)</strong>的回答,并提到'grad_norm 是先行指标'这一细节。
⚠️ Common Interview Pitfalls
  • ✕
    遇到 spike 直接重启训练(丢失进度,未定位根因)
  • ✕
    只调 lr 不检查数据(根因常在数据侧)
🎯 Interviewer Follow-ups
  • ?
    为什么 loss spike 后模型有时能自行恢复?
  • ?
    如何用 grad_norm 定位到具体的坏批次?
📚

Associated Knowledge Base Guides & Mindmaps

Explore the comprehensive technical article, exam cards, and global architecture tree.

← PreviousM3-059: Gradient Vanishing & Explosion: 解释残差连接与归一化如何共同保证梯度流。📋Back to BankNext →M3-061: Gradient Vanishing & Explosion: 解释 embedding 与输出层的梯度尺度问题(logit 缩放)。