M1-014M1: Mathematics & Statistics FundamentalsInformation TheoryMedium
Mastery:
Information Theory: 交叉熵损失与最大似然估计为什么等价?
📐 Mathematical Definition
⚡ Executive Summary
Core Concept: 最小化交叉熵 = 最大化对数似然;两者仅差一个与参数无关的常数(数据熵)。
📌 Key Takeaways
- •解释回归用 MSE 是高斯似然假设下的 MLE
- •分类用 CE 是 Categorical 似然下的 MLE
📐 Mathematical Derivations
推导只需一步:交叉熵 H(p,q_θ)=−Σp(x)log q_θ(x)=H(p)+KL(p‖q_θ),其中 H(p) 与 θ 无关,故 argmin_θ H(p,q_θ)=argmin_θ KL(p‖q_θ)。而在经验分布 p̂ 下,−Σp̂(x)log q_θ(x)=−(1/N)Σᵢlog q_θ(xᵢ) 正是负对数似然。所以'最小化交叉熵'与'最大化似然'是同一件事的两种叙述。更进一步,把 q_θ 换成不同分布假设即得不同损失:假设 y|x ~ N(f_θ(x),σ²) 则负对数似然正比于 ‖y−f_θ(x)‖²(MSE);假设 y|x ~ Laplace 则正比于 |y−f_θ(x)|(MAE);假设 Categorical 则得交叉熵(分类)。
🏭 Production Trade-offs
这个统一视角有两个重要的实践含义:① <strong>损失函数的选择 = 噪声分布的假设</strong>——所以当数据存在重尾噪声时 MSE 会被极端值主导,改用 Huber 或 MAE 更合理(对应更重的尾部假设);当预测计数时应用 Poisson 损失而非 MSE。② <strong>MSE 与 MAE 的梯度特性差异</strong>源于分布假设——MSE 梯度与误差成正比(对大误差敏感),MAE 梯度为常数符号(对异常值鲁棒但在 0 附近不平滑,需次梯度或 Huber 平滑)。此外,交叉熵在分类中与 softmax 组合时梯度恰好化简为 (p−y),非常干净,这也是它优于'MSE+softmax'的原因(后者梯度含 softmax 导数项,在饱和区极小)。
⚠️ Common Interview Pitfalls
- ✕认为 MSE 是'默认'损失而非某种分布假设的产物
- ✕对计数数据用 MSE(应使用 Poisson/NB 损失)
🎯 Interviewer Follow-ups
- ?MSE 对应什么分布假设?
- ?MAE 对应什么?(Laplace)
📚
Associated Knowledge Base Guides & Mindmaps
Explore the comprehensive technical article, exam cards, and global architecture tree.