M2-099M2: Classical Machine LearningRegularization (L1 / L2)Hard
Mastery:
Regularization (L1 / L2): 解释随机深度(Stochastic Depth / DropPath)与它与 Dropout 的差异。
📐 Mathematical Definition
⚡ Executive Summary
Core Concept: 训练时随机丢弃整个残差分支;推理时按概率缩放,与 BN 兼容性优于 dropout。
📌 Key Takeaways
- •丢弃的是整个层/分支(而非单个激活)
- •概率常随深度线性递增(浅层少丢、深层多丢)
📐 Mathematical Derivations
Stochastic Depth(Huang et al. 2016,用于 ResNet)的做法:训练时以概率 p 随机<strong>跳过整个残差分支</strong>(令 F_l 的输出为 0),只保留恒等路径 x_{l+1}=x_l;推理时使用全部分支并按 1/(1−p) 缩放(与 inverted dropout 一致)。<strong>与 Dropout 的差异</strong>:① <strong>粒度</strong>——dropout 丢弃<strong>单个神经元/激活</strong>,随机深度丢弃<strong>整个层/分支</strong>;② <strong>正则机制</strong>——dropout 通过破坏共适应与集成视角起正则作用,随机深度通过'随机深度网络集成'(每步是不同深度的子网络)起正则作用,且<strong>直接缩短了有效深度</strong>(缓解深层网络的优化与梯度问题);③ <strong>与 BN 的兼容性</strong>——这是关键差异:dropout 会改变激活的方差(除以 1−p),与 BN 的 running 统计冲突(方差偏移);而随机深度<strong>完全跳过分支</strong>(不改变保留分支的数值分布),故与 BN <strong>天然兼容</strong>——这使它成为 ResNet/ViT 等含 BN/LN 的深层网络的首选正则化手段。<strong>概率调度</strong>:常用<strong>线性递增</strong>——浅层 p 小(如 0.0)、深层 p 大(如 0.1–0.5),因为深层更冗余、丢弃影响小;也可用固定 p(如 0.1)。
🏭 Production Trade-offs
实践要点:① <strong>为什么深层网络特别需要</strong>——随着深度增加,层的冗余性上升(ResNet 的恒等路径使深层可被跳过),故随机深度的收益随深度增加;对 1000 层以上的网络收益显著。② <strong>与 DropPath 的关系</strong>——在 ViT/Transformer 中,同样的技术称为 <strong>DropPath</strong>(丢弃整个残差分支,作用于每个 token 的样本维度),是 ViT、Swin、ConvNeXt 训练的标准配置(常用 p=0.1–0.2)。③ <strong>推理期的实现</strong>——必须在 <code>eval</code> 模式下关闭丢弃并使用缩放因子;若忘记缩放会导致输出尺度偏小(这是常见 bug)。④ <strong>与 dropout 的取舍</strong>——现代深层网络(含 BN/LN)优先用 DropPath/随机深度;若必须用 dropout,应放在<strong>残差分支内</strong>(避免与 BN 冲突)。⑤ <strong>与其他正则的叠加</strong>——随机深度与 weight decay、数据增强、label smoothing 可叠加;但需注意总正则强度(过多正则导致欠拟合)。⑥ <strong>可解释性视角</strong>——随机深度使网络成为'深度随机集成',推理时的期望可视为对 2^L 个子网络的集成(类似 dropout 的集成解释)。
⚠️ Common Interview Pitfalls
- ✕把随机深度与 dropout 混用而不考虑 BN 兼容性
- ✕推理期忘记按 1/(1−p) 缩放
🎯 Interviewer Follow-ups
- ?为什么它与 BN 兼容?
- ?与 dropout 的机制差异?
📚
Associated Knowledge Base Guides & Mindmaps
Explore the comprehensive technical article, exam cards, and global architecture tree.