M1-066M1: Mathematics & Statistics FundamentalsInformation TheoryHard
Mastery:
Information Theory: 什么是 Wasserstein 距离?为什么 GAN 用它?
📐 Mathematical Definition
⚡ Executive Summary
Core Concept: 推土机距离:把分布 p 搬到 q 的最小代价;即使支撑不重叠也有连续梯度。
📌 Key Takeaways
- •满足度量公理(对称、三角不等式)
- •对支撑不重叠的分布仍提供梯度
- •Kantorovich-Rubinstein 对偶:W=sup_{‖f‖_L≤1}E_p[f]−E_q[f]
📐 Mathematical Derivations
Wasserstein 距离(又称 Earth Mover's Distance)定义为:在所有'把 p 的质输运到 q'的<strong>联合分布 γ</strong>(边缘分别为 p 和 q)中,最小化输运代价 E_{γ}[‖x−y‖]。直观上它衡量'把一堆土(p)搬成另一堆(q)所需的最小功'。<strong>与 KL/JS 的关键差异</strong>:① <strong>几何敏感性</strong>——当两个分布不重叠时,KL/JS 都是常数(无梯度),而 W 随分布间距<strong>线性变化</strong>(把土搬得越远代价越大),故提供连续的、有意义的梯度;② <strong>度量性质</strong>——W 满足对称性与三角不等式(是真正的度量),KL 不满足;③ <strong>弱拓扑下的收敛</strong>——W 收敛等价于分布弱收敛(含矩收敛),比 KL 的收敛要求更弱更实用。<strong>计算的对偶形式</strong>(Kantorovich-Rubinstein):W=sup_{‖f‖_L≤1} (E_p[f]−E_q[f]),即对所有 1-Lipschitz 函数的期望差取上确界。
🏭 Production Trade-offs
WGAN 的实现与要点:① <strong>对偶形式的可用性</strong>——用神经网络 f_θ(称为 critic,而非 discriminator)最大化 E_p[f]−E_q[f],同时约束 f 是 1-Lipschitz;损失为 −E_p[f]+E_q[f](没有 log、没有 sigmoid)。② <strong>Lipschitz 约束的近似</strong>——原始 WGAN 用 <strong>weight clipping</strong>(把权重裁到 [−c,c]),但会导致参数集中在边界、训练不稳;<strong>WGAN-GP</strong> 改用<strong>梯度惩罚</strong> λE[(‖∇_x f(x̂)‖₂−1)²](x̂ 是真实与生成样本间的随机插值点),效果好但计算贵;<strong>谱归一化(spectral normalization)</strong> 用最大奇异值约束每层,是更高效的替代(SN-GAN)。③ <strong>实际收益</strong>——WGAN 的 critic 损失与生成质量<strong>相关</strong>(可作训练监控指标),而原始 GAN 的 discriminator 损失不相关;训练更稳定、模式崩溃更少。④ <strong>局限</strong>——高维下 W 的估计困难(对偶形式需搜索所有 Lipschitz 函数)、Lipschitz 约束的近似引入偏差、计算成本高于原始 GAN。
⚠️ Common Interview Pitfalls
- ✕把 weight clipping 当作 Lipschitz 约束的最佳实现(WGAN-GP 更好)
- ✕认为 W 可以精确计算(高维下只能近似)
🎯 Interviewer Follow-ups
- ?为什么它比 KL/JS 更适合 GAN?
- ?如何近似计算它?(WGAN-GP)
📚
Associated Knowledge Base Guides & Mindmaps
Explore the comprehensive technical article, exam cards, and global architecture tree.