M2-013M2: Classical Machine LearningRegularization (L1 / L2)Medium
Mastery:

Regularization (L1 / L2): Dropout 为什么能起到正则化作用?给出两种解释。

📐 Mathematical Definition
train:h′=m⊙h1−p,eval:h′=h\text{train}: h'=\frac{m\odot h}{1-p},\qquad \text{eval}: h'=h
⚡ Executive Summary
Core Concept: ① 集成视角:等价于指数级子网络的 bagging;② 破坏共适应,迫使特征独立有用。

📌 Key Takeaways

  • •
    推理期需缩放(inverted dropout)
  • •
    与 BN 同用需注意方差偏移

📐 Mathematical Derivations

两种解释:① <strong>集成视角</strong>(Srivastava et al. 2014)——每次前向随机丢弃一部分神经元,相当于训练了 2ⁿ 个共享权重的子网络;推理时用全部神经元并缩放,相当于对这些子网络的<strong>几何平均</strong>做集成。集成降方差,故有正则效果。② <strong>共适应破坏视角</strong>——神经元不能依赖特定同伴的存在(因为同伴可能被丢弃),被迫学习<strong>独立有用</strong>的特征。这抑制了'特征间互相补偿'的脆弱依赖,提升鲁棒性。<strong>Inverted dropout</strong> 的实现细节:训练时对保留的激活除以 (1−p)(放大),推理时恒等——这样训练与推理的期望一致,避免推理期额外缩放。

🏭 Production Trade-offs

关键权衡与坑:① <strong>p 的选择</strong>——输入层通常 p=0.1–0.2,隐藏层 0.3–0.5;p 过大导致欠拟合与训练不稳,p 过小无效果。② <strong>与 BatchNorm 的冲突</strong>——BN 在训练时用 batch 统计、推理时用 running 统计,而 dropout 改变激活的方差,导致训练/推理的统计不一致(方差偏移,variance shift);常见解法是把 dropout 放在残差分支上(而非主路径)、或用 LN/GN 替代 BN、或降低 p。③ <strong>现代趋势</strong>——大模型预训练常用 dropout=0(数据量大时正则需求低),而微调时用 0.1;Transformer 中 dropout 常用于注意力权重与残差连接处。④ <strong>与 weight decay 的关系</strong>——两者机制不同(dropout 作用于激活、weight decay 作用于参数),可叠加;但 dropout 的隐式正则已较强时,需相应减小 weight decay。
⚠️ Common Interview Pitfalls
  • ✕
    推理时忘记关闭 dropout(输出随机、期望被放大)
  • ✕
    dropout 与 BN 直接叠加而不处理方差偏移
🎯 Interviewer Follow-ups
  • ?
    为什么推理时不能开 dropout?
  • ?
    Dropout 与 weight decay 的关系?
📚

Associated Knowledge Base Guides & Mindmaps

Explore the comprehensive technical article, exam cards, and global architecture tree.

← PreviousM2-012: Regularization (L1 / L2): 正则化系数 λ 如何选择?过大过小分别会怎样。📋Back to BankNext →M2-014: Regularization (L1 / L2): BatchNorm 为什么也有正则化效果?它和 Dropout 能一起用吗。