M2-014M2: Classical Machine LearningRegularization (L1 / L2)Medium
Mastery:

Regularization (L1 / L2): BatchNorm 为什么也有正则化效果?它和 Dropout 能一起用吗。

📐 Mathematical Definition
x^=x−μBσB2+ϵ\hat x=\frac{x-\mu_B}{\sqrt{\sigma_B^2+\epsilon}}
⚡ Executive Summary
Core Concept: BN 引入 mini-batch 统计噪声,起轻微正则作用;但与 Dropout 同用会互相干扰(方差偏移)。

📌 Key Takeaways

  • •
    BN 主要目的是稳定优化,正则是副产品
  • •
    替代方案:LN/GN 或把 dropout 放残差分支

📐 Mathematical Derivations

BN 的正则化机制:训练时每个 batch 的均值/方差是<strong>有噪声的估计</strong>(尤其小 batch),这给激活注入了随机性,效果类似 dropout——网络不能依赖精确的激活值。因此 BN 的原始论文观察到'使用 BN 后可以减小或去掉 dropout'。但必须强调:<strong>BN 的主要目的是稳定优化</strong>(缓解内部协变量偏移、允许更大学习率、平滑损失景观),正则化只是副产品,且在小 batch 下噪声过大反而有害。

🏭 Production Trade-offs

与 Dropout 同用的冲突与解法:① <strong>方差偏移问题</strong>——dropout 改变激活方差(除以 1−p),而 BN 的 running 统计是在<strong>有 dropout 的情况下</strong>估计的;推理时 dropout 关闭使方差变化,导致 BN 的归一化失配,性能下降。② <strong>常见解法</strong>——(a) 把 dropout 放在<strong>残差分支</strong>(BN 之前或之后分离);(b) 用 <strong>LN/GN</strong> 替代 BN(它们不依赖 batch 统计,训练/推理一致);(c) 减小 dropout 的 p(如 0.1);(d) 使用 <strong>DropPath / Stochastic Depth</strong>(按样本丢弃整个残差分支,与 BN 兼容)。③ <strong>现代实践</strong>——Transformer 用 LN 故无此问题;CNN 中 ResNet 系列把 BN 放在卷积后、dropout 尽量少用或只在全连接层;检测/分割常用 GN。
⚠️ Common Interview Pitfalls
  • ✕
    认为 BN 的正则效果可以替代 dropout(机制不同,小 batch 下更糟)
  • ✕
    在 BN 后直接加 dropout 而不考虑方差偏移
🎯 Interviewer Follow-ups
  • ?
    BN 在推理期为什么用 running 统计?
  • ?
    为什么 BN 对小 batch 效果差?
📚

Associated Knowledge Base Guides & Mindmaps

Explore the comprehensive technical article, exam cards, and global architecture tree.

← PreviousM2-013: Regularization (L1 / L2): Dropout 为什么能起到正则化作用?给出两种解释。📋Back to BankNext →M2-015: Regularization (L1 / L2): 解释早停(early stopping)为什么等价于某种正则化。