M1-012M1: Mathematics & Statistics FundamentalsInformation TheoryMedium
Mastery:

Information Theory: KL 散度为什么不对称?前向与反向 KL 在优化上有什么区别?

📐 Mathematical Definition
KL(p∥q)=∑plog⁡pq,KL(q∥p)=∑qlog⁡qp\mathrm{KL}(p\|q)=\sum p\log\frac{p}{q},\qquad \mathrm{KL}(q\|p)=\sum q\log\frac{q}{p}
⚡ Executive Summary
Core Concept: KL(p‖q) ≠ KL(q‖p)。前向(moment-covering)逼 q 覆盖 p 的支撑;反向(mode-seeking)让 q 收缩到 p 的众数。

📌 Key Takeaways

  • •
    变分推断用反向 KL → 欠估计方差(mode seeking)
  • •
    最大似然等价于最小化前向 KL
  • •
    VAE 的 ELBO 含反向 KL,是后验近似收缩的来源

📐 Mathematical Derivations

不对称的根源在于 KL 是<strong>加权求和</strong>,而权重是第一个分布:KL(p‖q)=Σp·log(p/q),权重为 p;KL(q‖p) 权重为 q。因此两者的'惩罚重点'完全不同——<strong>前向 KL(p‖q)</strong> 在 p(x)>0 而 q(x)→0 时惩罚 →∞,迫使 q 覆盖 p 的整个支撑(zero-avoiding / mass-covering),但允许 q 在 p 很小的区域铺开;<strong>反向 KL(q‖p)</strong> 在 q(x)>0 而 p(x)→0 时惩罚 →∞,迫使 q 收缩到 p 的支撑内(zero-forcing / mode-seeking),但允许多个众数只保留一个。

🏭 Production Trade-offs

三种典型用法的对应:① <strong>最大似然估计 ≡ 最小化前向 KL(p_data‖p_model)</strong>——所以 MLE 会覆盖所有模式(这就是为什么用 MLE 训练的生成模型不会丢模式);② <strong>变分推断最小化反向 KL(q_approx‖p_posterior)</strong>——所以变分后验倾向于<strong>低估方差</strong>并聚焦单一模式,这是 VAE 生成偏模糊的理论根源(可通过 IWAE、β-VAE 缓解);③ <strong>GAN 的 JS 散度</strong>在两个分布支撑不重叠时饱和为常数 log2,梯度消失——这是 WGAN 改用 Wasserstein 距离(有连续梯度)的直接动机。对称化方案:JS 散度((KL(p‖m)+KL(q‖m))/2,m 为混合)、Wasserstein 距离、或 Maximum Mean Discrepancy(MMD)。
⚠️ Common Interview Pitfalls
  • ✕
    以为 KL 是距离(它不满足对称性与三角不等式)
  • ✕
    在变分推断中期望前向 KL 的行为(会得到 mode-seeking 的结果)
🎯 Interviewer Follow-ups
  • ?
    如何构造对称散度?(JS 散度 / Wasserstein)
  • ?
    GAN 为什么用 JS 会导致梯度消失?
📚

Associated Knowledge Base Guides & Mindmaps

Explore the comprehensive technical article, exam cards, and global architecture tree.

← PreviousM1-011: Information Theory: 定义信息量(自信息)、熵、交叉熵、KL 散度,并说明相互关系。📋Back to BankNext →M1-013: Information Theory: 解释互信息与点互信息(PMI),它们分别用在哪里?