M3-065M3: Deep Learning FoundationsLoss Functions & ObjectivesEasy
Mastery:

Loss Functions & Objectives: 解释交叉熵与 BCE 的关系与数值稳定写法。

📐 Mathematical Definition
CE=−∑iyilog⁡pi;BCEstable=max⁡(z,0)−z⋅y+log⁡(1+e−∣z∣)\text{CE}=-\sum_i y_i\log p_i;\qquad \text{BCE}_{\text{stable}}=\max(z,0)-z\cdot y+\log(1+e^{-|z|})
⚡ Executive Summary
Core Concept: BCE 是二分类下的交叉熵特例;实现上把 sigmoid 与 CE 合并为 logits 形式(log-sum-exp)以避免溢出。

📌 Key Takeaways

  • •
    多分类 CE = softmax + NLL;二分类 BCE = sigmoid + NLL
  • •
    直接算 log(sigmoid(z)) 在 |z| 大时溢出,需用 logits 形式
  • •
    PyTorch 的 cross_entropy_with_logits / BCEWithLogitsLoss 已内置稳定实现

📐 Mathematical Derivations

数学机理:<strong>交叉熵</strong> CE=−Σ y_i log p_i 度量'用预测分布 p 编码真实分布 y 的额外代价';当 y 为 one-hot 时 CE=−log p_c(正确类概率的负对数)。<strong>BCE</strong> 是 K=2 的 CE:BCE=−[y log p+(1−y)log(1−p)],其中 p=σ(z)。<strong>数值问题</strong>:若先算 p=σ(z) 再取 log,则 z 很大时 p→1、log p→0 尚可;但 z 很负时 p→0、log(p) 需要计算 log(0)=−inf,且 1−p 会因浮点舍入<strong>精确变成 0</strong>(p 下溢),导致 log(1−p)=−inf、loss=inf。<strong>稳定写法</strong>是把 sigmoid 与 log 合并:对 BCE,利用 σ 的性质可推出 <strong>BCE = max(z,0) − z·y + log(1+e^{−|z|})</strong>,其中 log(1+e^{−|z|}) 中的指数恒为负(≤0)、不会溢出;这一形式对所有 z 都数值稳定。<strong>多分类 CE</strong> 同理:softmax+CE 合并为 <strong>log-sum-exp</strong> 形式 loss = −z_c + log Σ_j e^{z_j},实现时先减去 max_j z_j(max-subtraction),使指数 ≤0。这正是 PyTorch 提供 <code>cross_entropy</code>(接受 logits)与 <code>BCEWithLogitsLoss</code> 的原因——<strong>不要</strong>手动先做 sigmoid/softmax 再算 CE。

🏭 Production Trade-offs

深度剖析与工程权衡:① <strong>log-sum-exp 是通用原语</strong>——它同时解决'上溢'(减最大值)与'下溢'(用 log1p 形式),是 softmax、CE、log-domain 概率运算的基础;面试中若被要求手写数值稳定的 softmax,应给出'减 max + log-sum-exp'的完整版本。② <strong>与 label smoothing 的交互</strong>——label smoothing 把目标从 one-hot 变为软标签,此时 CE 变为 −Σ ỹ log p,实现上仍是 log-softmax 的加权和,但<strong>不能</strong>用'−z_c + lse'的简化式(需完整加权)。③ <strong>类别数极多时的效率</strong>——LLM 的 V 可达 128k,全量 softmax 的 log-sum-exp 是主要开销;故有 <strong>sampled softmax / hierarchical softmax / 负采样</strong> 等近似,以及 <strong>flash-attention 式的在线 softmax</strong>(分块计算、在线更新 max 与 sum)。④ <strong>与 focal loss 的关系</strong>——focal loss 在 CE 上加调制因子 (1−p_t)^γ,实现上同样需稳定形式(用 logits 计算 p_t)。⑤ <strong>温度参数</strong>——logits 除以温度 T 后再算 CE,等价于软化/锐化分布;蒸馏中教师用高温、学生用 T=1。⑥ <strong>面试要点</strong>——被问'手写交叉熵',务必写出 log-sum-exp 稳定版并解释'为什么不能先 softmax 再 log';这是'工程细节意识'的直接体现。
⚠️ Common Interview Pitfalls
  • ✕
    先算 sigmoid/softmax 再取 log(数值溢出/下溢)
  • ✕
    在类别数大时忽略 log-sum-exp 的在线分块实现
🎯 Interviewer Follow-ups
  • ?
    为什么 PyTorch 不推荐先 sigmoid 再 BCE?
  • ?
    softmax 的 log-sum-exp 技巧如何实现?
📚

Associated Knowledge Base Guides & Mindmaps

Explore the comprehensive technical article, exam cards, and global architecture tree.

← PreviousM3-064: Loss Functions & Objectives: 比较 MSE、MAE、Huber 损失。📋Back to BankNext →M3-066: Loss Functions & Objectives: 解释 Focal Loss 的设计动机与公式。