M3-002M3: Deep Learning FoundationsBackprop & AutodiffEasy
Mastery:

Backprop & Autodiff: 反向传播需要保存哪些中间量?如何省显存。

📐 Mathematical Definition
checkpoint: recompute activations in backward\text{checkpoint}:\ \text{recompute activations in backward}
⚡ Executive Summary
Core Concept: 需保存激活值与局部梯度;可用梯度检查点、混合精度、分片(FSDP)降低显存。

📌 Key Takeaways

  • •
    检查点用计算换内存(约 √n 内存)
  • •
    FSDP/ZeRO 分片优化器状态与梯度

📐 Mathematical Derivations

保存内容:反向传播需要 (a) <strong>前向的中间激活</strong>(用于计算局部雅可比,如 attention 的 softmax 输出、激活函数输出);(b) <strong>局部导数所需的额外量</strong>(如 BatchNorm 的均值方差、dropout 的 mask);(c) 参数本身(若参数被覆盖则需保留副本)。<strong>显存分布</strong>(Adam + FP32 主权重 + FP16 计算,混合精度):参数 2B、梯度 2B、FP32 主权重 4B、Adam 的 m 与 v 各 4B → <strong>优化器状态与主权重合计 16B/参数</strong>(远超参数本身);而<strong>激活</strong>与 batch×序列长度成正比(长序列时成为主要瓶颈)。<strong>四类省显存手段</strong>:① <strong>梯度检查点</strong>——只保存少数'检查点'的激活,反向时重算中间激活,显存从 O(n) 降到 O(√n),代价约 30% 额外计算;② <strong>混合精度</strong>——激活用 FP16/BF16(减半);③ <strong>分片</strong>——FSDP/ZeRO 把优化器状态、梯度、参数分片到多卡;④ <strong>激活卸载(offloading)</strong>——把激活换出到 CPU 内存(用 PCIe 带宽换显存)。

🏭 Production Trade-offs

实践要点:① <strong>检查点的时间-内存权衡</strong>——√n 的来源是'分块检查点':把 L 层分成 √L 块、每块存一个检查点,反向时逐块重算;时间代价约为 1 次额外前向(30% 左右)。② <strong>哪些层值得检查</strong>——激活大的层(attention、大 FFN)收益最大;小层(LayerNorm)不值得。③ <strong>与混合精度的叠加</strong>——两者独立可叠加;但注意 FP16 激活下重算需保持数值一致(框架自动处理)。④ <strong>优化器状态的优化</strong>——<strong>8-bit Adam</strong>(bitsandbytes)把 m、v 量化为 8-bit,显存从 8B 降到 2B/参数;<strong>Adafactor</strong> 用分解近似二阶矩(只存行/列统计),显存从 O(n) 降到 O(√n)。⑤ <strong>长序列的特殊问题</strong>——激活 ∝ batch×seq×hidden×layers,长上下文时激活常超过参数与优化器状态;此时应优先用检查点 + FlashAttention(不物化 L×L 矩阵)。⑥ <strong>诊断</strong>——用 <code>torch.cuda.max_memory_allocated()</code> 与显存剖析工具确认瓶颈在激活还是优化器状态,再决定优化方向(这是省显存的第一步)。
⚠️ Common Interview Pitfalls
  • ✕
    盲目开检查点而不诊断瓶颈(可能优化错方向)
  • ✕
    忽略优化器状态占 16B/参数这一事实
🎯 Interviewer Follow-ups
  • ?
    检查点的时间代价?
  • ?
    为什么优化器状态占显存最多?
📚

Associated Knowledge Base Guides & Mindmaps

Explore the comprehensive technical article, exam cards, and global architecture tree.

← PreviousM3-001: Backprop & Autodiff: 解释计算图与自动微分的两种模式(前向/反向)。📋Back to BankNext →M3-003: Backprop & Autodiff: 解释梯度累加(gradient accumulation),它等价于什么。