M3-023M3: Deep Learning FoundationsNormalization TechniquesMedium
Mastery:
Normalization Techniques: 解释 BatchNorm 在推理时为什么用 running 统计。
📐 Mathematical Definition
⚡ Executive Summary
Core Concept: 推理时无 batch(或 batch=1),且需确定性输出;running 统计是训练期 batch 统计的滑动平均。
📌 Key Takeaways
- •train/eval 行为不一致是常见 bug 源
- •LN/GN 无此问题
📐 Mathematical Derivations
两个原因:① <strong>推理时可能没有 batch</strong>——线上推理常是单样本或小 batch(batch=1 时 batch 方差为 0,无法归一化);② <strong>确定性要求</strong>——推理结果不应依赖'同批处理的其他样本'(否则同一输入在不同 batch 组合下输出不同,不可复现、且引入信息泄漏)。<strong>running 统计的构造</strong>:训练时每个 batch 计算 μ_B、σ²_B,并用滑动平均更新全局估计:μ_run←m·μ_run+(1−m)·μ_B(m 通常 0.9–0.99,等价于对最近约 1/(1−m) 个 batch 的指数加权平均);σ² 同理。推理时用 μ_run、σ²_run 归一化,<strong>不再更新</strong>。<strong>为什么这样合理</strong>:训练结束后 μ_run 近似了'整个训练集的平均统计量',故推理时用它是无偏且确定的;这与'用训练集统计量做标准化'的直觉一致。<strong>常见 bug</strong>:① <strong>忘记 <code>model.eval()</code></strong>——推理时仍用 batch 统计,导致结果依赖 batch 组成、且随 batch 大小变化;② <strong>running 统计未收敛</strong>——训练步数太少(或 momentum 太小)时 μ_run 仍偏离真实分布,推理性能差;③ <strong>微调时冻结 BN 统计</strong>——若微调数据分布与预训练不同,应更新 running 统计(或改用 GN/LN),否则归一化失配。
🏭 Production Trade-offs
实践要点:① <strong>微调 BN 的策略</strong>——(a) 数据分布相似、batch 大 → 正常更新 running 统计;(b) 数据分布不同或 batch 小 → <strong>冻结 BN 统计</strong>(只训练 γ、β)或改用 GN/LN;这是迁移学习中的常见决策点。② <strong>与 SyncBN</strong>——多卡训练时若各卡 batch 小,可用 SyncBN(跨卡同步统计量,等效于大 batch);代价是通信开销。③ <strong>BN 与 dropout 的冲突</strong>——dropout 改变激活方差,而 BN 的 running 统计是在有 dropout 时估计的,推理时 dropout 关闭导致方差偏移(见'BN 与 dropout'题)。④ <strong>BN 的初始化</strong>——γ=1、β=0(恒等);running_mean=0、running_var=1(初始假设标准正态)。⑤ <strong>诊断方法</strong>——对比训练与推理模式下的输出(同一输入),若差异大说明 running 统计有问题;也可打印 μ_run 与训练集真实均值对比。⑥ <strong>替代方案</strong>——若 batch 太小或推理需确定性,用 LN/GN(它们训练推理一致,无此问题);这也是 Transformer 用 LN 的原因之一。
⚠️ Common Interview Pitfalls
- ✕推理时未切换到 eval 模式(BN 用 batch 统计)
- ✕微调时盲目冻结/更新 BN 统计而不检查分布差异
🎯 Interviewer Follow-ups
- ?为什么 BN 在推理时不能用当前 batch 统计?
- ?BN 的常见 bug 有哪些?
📚
Associated Knowledge Base Guides & Mindmaps
Explore the comprehensive technical article, exam cards, and global architecture tree.