M1-015M1: Mathematics & Statistics FundamentalsInformation TheoryHard
Mastery:
Information Theory: 解释困惑度(Perplexity)的物理意义,为什么它等于 exp(交叉熵)。
📐 Mathematical Definition
⚡ Executive Summary
Core Concept: PPL 是模型在每步'等效犹豫的候选数';PPL=exp(CE),越低越好。
📌 Key Takeaways
- •PPL 只在同一 tokenizer 下可比
- •PPL 低不代表下游任务好(目标错配)
📐 Mathematical Derivations
困惑度的定义是<strong>平均负对数似然的指数</strong>:PPL=exp(CE)。它的直观解释来自均匀分布——若模型对每个 token 都在 V 个候选上均匀分布,则 CE=log V,PPL=V,即'模型每步等效地在 V 个选项中犹豫'。因此 PPL=10 意味着模型每步的不确定性相当于从 10 个等概率候选中选一个。数学上它等于<strong>模型分布下的平均分支因子</strong>,也是算术平均的倒数几何平均:PPL=(Πᵗ1/p(x_t))^{1/N}。信息论上,CE 的单位是 nat(自然对数)或 bit(以 2 为底),PPL 是它的指数化,因此<strong>PPL 的单调变换不改变模型排序</strong>,比较模型时用 CE 或 PPL 等价。
🏭 Production Trade-offs
两个必须注意的陷阱:① <strong>跨 tokenizer 不可比</strong>——PPL 是 per-token 的,若模型 A 的词表更大(每 token 承载更多信息),它的 PPL 天然更低,这并非模型更强。公平比较应换算到 <strong>bits-per-byte(BPB)</strong>,因为字节是语言无关的单位:BPB=CE/(ln2 × 平均字节数/token)。这也是为什么 LLaMA 系列论文同时报告 PPL 与 BPB。② <strong>PPL 与生成质量弱相关</strong>——PPL 衡量的是对训练分布的拟合,不衡量指令遵循、事实性、有用性;一个 PPL 极低的模型可能因为训练数据分布偏窄而在开放任务上表现差。此外,PPL 只在同分布验证集上有意义,跨领域会失真。
⚠️ Common Interview Pitfalls
- ✕跨 tokenizer 直接比较 PPL
- ✕把 PPL 当作生成质量或指令遵循能力的指标
🎯 Interviewer Follow-ups
- ?跨 tokenizer 如何比较模型?
- ?为什么 PPL 与生成质量弱相关?
📚
Associated Knowledge Base Guides & Mindmaps
Explore the comprehensive technical article, exam cards, and global architecture tree.