M3-053M3: Deep Learning FoundationsRegularization & Training TricksHard
Mastery:

Regularization & Training Tricks: 解释 dropout 在 Transformer / LLM 中的使用差异。

📐 Mathematical Definition
dropout: h′=h⊙m/(1−p), m∼Bernoulli(1−p)\text{dropout}:\ h'=h\odot m/(1-p),\ m\sim\mathrm{Bernoulli}(1-p)
⚡ Executive Summary
Core Concept: 小模型/CV/微调用 dropout(0.1);大 LLM 预训练几乎不用(数据充足、显存开销大、与 LN 功能重叠)。

📌 Key Takeaways

  • •
    dropout 等价于子网络集成 + 隐式 L2 正则
  • •
    LLM 预训练通常只保留 attention dropout=0
  • •
    大模型靠数据规模与 weight decay 正则,而非 dropout

📐 Mathematical Derivations

数学机理:<strong>dropout</strong> 在训练时以概率 p 随机置零激活、并以 1/(1−p) 缩放(inverted dropout,保证期望不变);推理时不丢弃。理论解释有二:<strong>(1) 集成视角</strong>——训练时相当于采样 2ⁿ 个子网络(n 为神经元数),推理时用全部神经元(近似子网络集成);<strong>(2) 正则视角</strong>——阻止神经元间'共适应'(co-adaptation),迫使每个单元独立有用。<strong>为什么大 LLM 不用</strong>:(a) <strong>数据充足</strong>——LLM 预训练数据量远超参数量,过拟合不是主要矛盾,正则收益低;(b) <strong>成本</strong>——dropout 需要额外随机数与掩码、破坏 kernel 融合、降低 GPU 利用率,在超大规模训练中代价显著;(c) <strong>功能重叠</strong>——LayerNorm/残差连接已经提供了稳定训练的作用,dropout 的边际收益下降;(d) <strong>与 MoE/长序列的冲突</strong>——随机丢弃会加剧训练-推理不一致。实证上,GPT-3/LLaMA 等仅保留极少量 dropout(常为 0)。

🏭 Production Trade-offs

深度剖析与工程权衡:① <strong>微调场景反转</strong>——在<strong>小数据微调</strong>(如几千条指令)时 dropout 仍有效,故很多 SFT 配置用 dropout=0.1;但 QLoRA 等高效微调常设 0(因为 LoRA 本身的低秩结构已提供正则)。② <strong>注意力 dropout</strong>——部分模型保留 attention 权重的 dropout(防止注意力过度集中),但主流大模型也设为 0。③ <strong>Stochastic depth</strong>——ViT 等用随机跳过整层(drop path),是 dropout 的层级版本,在大视觉模型(ViT-H)中仍有效;LLM 中较少用。④ <strong>与 weight decay 的分工</strong>——大模型的正则主要交给 weight decay 与'早停/数据规模';dropout 退居次要。⑤ <strong>历史教训</strong>——Transformer 原论文用 dropout=0.1;随着数据与规模增长,dropout 逐渐被移除(BERT 0.1 → GPT-3 0.0),这本身是'正则强度随数据规模下降'的经典案例。⑥ <strong>面试要点</strong>——被问'为什么大模型不用 dropout',应从<strong>数据规模、显存/算力成本、与 LN 的功能重叠</strong>三个角度回答,而非'因为大模型不需要正则'这种笼统说法。
⚠️ Common Interview Pitfalls
  • ✕
    在大规模预训练上保留高 dropout(浪费算力且无收益)
  • ✕
    忽略 dropout 会破坏 kernel 融合、降低吞吐
🎯 Interviewer Follow-ups
  • ?
    为什么大模型不需要 dropout?
  • ?
    dropout 与 LayerNorm 的功能重叠在哪里?
📚

Associated Knowledge Base Guides & Mindmaps

Explore the comprehensive technical article, exam cards, and global architecture tree.

← PreviousM3-052: Regularization & Training Tricks: 解释 label smoothing 的作用与代价。📋Back to BankNext →M3-054: Gradient Vanishing & Explosion: 解释梯度消失与爆炸的成因与缓解。