M3-053M3: Deep Learning FoundationsRegularization & Training TricksHard
Mastery:
Regularization & Training Tricks: 解释 dropout 在 Transformer / LLM 中的使用差异。
📐 Mathematical Definition
⚡ Executive Summary
Core Concept: 小模型/CV/微调用 dropout(0.1);大 LLM 预训练几乎不用(数据充足、显存开销大、与 LN 功能重叠)。
📌 Key Takeaways
- •dropout 等价于子网络集成 + 隐式 L2 正则
- •LLM 预训练通常只保留 attention dropout=0
- •大模型靠数据规模与 weight decay 正则,而非 dropout
📐 Mathematical Derivations
数学机理:<strong>dropout</strong> 在训练时以概率 p 随机置零激活、并以 1/(1−p) 缩放(inverted dropout,保证期望不变);推理时不丢弃。理论解释有二:<strong>(1) 集成视角</strong>——训练时相当于采样 2ⁿ 个子网络(n 为神经元数),推理时用全部神经元(近似子网络集成);<strong>(2) 正则视角</strong>——阻止神经元间'共适应'(co-adaptation),迫使每个单元独立有用。<strong>为什么大 LLM 不用</strong>:(a) <strong>数据充足</strong>——LLM 预训练数据量远超参数量,过拟合不是主要矛盾,正则收益低;(b) <strong>成本</strong>——dropout 需要额外随机数与掩码、破坏 kernel 融合、降低 GPU 利用率,在超大规模训练中代价显著;(c) <strong>功能重叠</strong>——LayerNorm/残差连接已经提供了稳定训练的作用,dropout 的边际收益下降;(d) <strong>与 MoE/长序列的冲突</strong>——随机丢弃会加剧训练-推理不一致。实证上,GPT-3/LLaMA 等仅保留极少量 dropout(常为 0)。
🏭 Production Trade-offs
深度剖析与工程权衡:① <strong>微调场景反转</strong>——在<strong>小数据微调</strong>(如几千条指令)时 dropout 仍有效,故很多 SFT 配置用 dropout=0.1;但 QLoRA 等高效微调常设 0(因为 LoRA 本身的低秩结构已提供正则)。② <strong>注意力 dropout</strong>——部分模型保留 attention 权重的 dropout(防止注意力过度集中),但主流大模型也设为 0。③ <strong>Stochastic depth</strong>——ViT 等用随机跳过整层(drop path),是 dropout 的层级版本,在大视觉模型(ViT-H)中仍有效;LLM 中较少用。④ <strong>与 weight decay 的分工</strong>——大模型的正则主要交给 weight decay 与'早停/数据规模';dropout 退居次要。⑤ <strong>历史教训</strong>——Transformer 原论文用 dropout=0.1;随着数据与规模增长,dropout 逐渐被移除(BERT 0.1 → GPT-3 0.0),这本身是'正则强度随数据规模下降'的经典案例。⑥ <strong>面试要点</strong>——被问'为什么大模型不用 dropout',应从<strong>数据规模、显存/算力成本、与 LN 的功能重叠</strong>三个角度回答,而非'因为大模型不需要正则'这种笼统说法。
⚠️ Common Interview Pitfalls
- ✕在大规模预训练上保留高 dropout(浪费算力且无收益)
- ✕忽略 dropout 会破坏 kernel 融合、降低吞吐
🎯 Interviewer Follow-ups
- ?为什么大模型不需要 dropout?
- ?dropout 与 LayerNorm 的功能重叠在哪里?
📚
Associated Knowledge Base Guides & Mindmaps
Explore the comprehensive technical article, exam cards, and global architecture tree.