M3-052M3: Deep Learning FoundationsRegularization & Training TricksMedium
Mastery:
Regularization & Training Tricks: 解释 label smoothing 的作用与代价。
📐 Mathematical Definition
⚡ Executive Summary
Core Concept: 把硬标签替换为软标签(正确类 1−ε、其余 ε/(K−1));抑制过度自信、改善校准,但损失可解释性与蒸馏能力。
📌 Key Takeaways
- •ε 常取 0.1;等价于对 logits 加正则
- •提升校准与鲁棒性,但降低'知识蒸馏'效果
- •会使模型输出有界(logit 差距不会无限大)
📐 Mathematical Derivations
数学机理:标准交叉熵鼓励正确类 logit 趋于无穷(因为 loss 无下界、只要 logit 差距足够大 loss 就趋于 0),导致<strong>过度自信</strong>与过大的 logit 幅度。<strong>Label smoothing</strong> 把目标从 one-hot 变为'正确类 1−ε、其余类各 ε/(K−1)'。在最优解处,交叉熵的梯度为零要求模型输出等于平滑后的标签,即 p_c=1−ε、其余 ε/(K−1)——<strong>logit 的差距被限制在有限值</strong>(正比于 log((1−ε)(K−1)/ε))。这带来三个效果:(1) <strong>校准改善</strong>——置信度更接近真实准确率;(2) <strong>鲁棒性提升</strong>——对标签噪声与对抗样本更鲁棒(不依赖极端 logit);(3) <strong>泛化改善</strong>——在 ImageNet 等任务上稳定提升。<strong>代价</strong>:(a) <strong>蒸馏失效</strong>——Hinton 的蒸馏依赖教师模型的软标签携带'类间相似性'信息(如'3'与'8'相似),而 label smoothing 是<strong>均匀分布</strong>、不含类间结构,故教师被平滑后蒸馏效果下降(Müller 等 2019 明确指出);(b) <strong>置信度不可用</strong>——若下游需要真实置信度做拒识/校准,平滑后的输出被系统性压低。
🏭 Production Trade-offs
深度剖析与工程权衡:① <strong>与温度缩放的互补</strong>——label smoothing 在<strong>训练时</strong>限制 logit 幅度,温度缩放在<strong>推理时</strong>调整置信度;两者都改善校准但机制不同。② <strong>ε 的选择</strong>——ε=0.1 是 ImageNet 的常用值;过大(如 0.3)会欠拟合(目标过于模糊)、损害精度。③ <strong>在 LLM 中的使用</strong>——LLM 预训练常<strong>不用</strong> label smoothing(因 next-token 分布本身有信息、且需保留概率结构);但某些翻译/分类微调任务会用。④ <strong>与 Mixup 的区别</strong>——label smoothing 只软化标签(样本不变),Mixup 同时软化样本与标签;前者更便宜。⑤ <strong>'知识蒸馏被破坏'的修复</strong>——若既要平滑又要蒸馏,可用'softened 但保留结构的标签'(如对教师分布做小温度缩放)而非均匀平滑。⑥ <strong>面试要点</strong>——加分回答是主动指出'label smoothing 的代价:破坏类间相似性信息、损害蒸馏';这是 Müller 等 2019 《When Does Label Smoothing Help?》的核心结论,能显著体现文献掌握。
⚠️ Common Interview Pitfalls
- ✕以为 label smoothing 只有好处(会损害蒸馏与类间结构)
- ✕在下游需要真实置信度时仍用强平滑
🎯 Interviewer Follow-ups
- ?为什么 label smoothing 会损害知识蒸馏?
- ?label smoothing 与温度缩放的关系?
📚
Associated Knowledge Base Guides & Mindmaps
Explore the comprehensive technical article, exam cards, and global architecture tree.