M1-074M1: Mathematics & Statistics Fundamentals估计理论 (MLE/MAP)Medium
Mastery:

估计理论 (MLE/MAP): 解释 Fisher 信息量与 Cramér-Rao 下界。

📐 Mathematical Definition
I(θ)=E[(∂log⁡p∂θ)2]=−E[∂2log⁡p∂θ2],Var(θ^)≥1nI(θ)I(\theta)=\mathbb E\Big[\Big(\frac{\partial\log p}{\partial\theta}\Big)^2\Big]=-\mathbb E\Big[\frac{\partial^2\log p}{\partial\theta^2}\Big],\qquad \mathrm{Var}(\hat\theta)\ge\frac{1}{nI(\theta)}
⚡ Executive Summary
Core Concept: Fisher 信息是对数似然关于参数的曲率(二阶导期望);CRB 给出无偏估计方差的下界。

📌 Key Takeaways

  • •
    Fisher 信息可加(独立样本相加)
  • •
    达到 CRB 的估计量称为有效估计量(MLE 渐近有效)

📐 Mathematical Derivations

Fisher 信息量的两种等价定义:① 得分函数平方的期望 I(θ)=E[(∂log p/∂θ)²];② 对数似然二阶导的负期望 I(θ)=−E[∂²log p/∂θ²]。第二种形式给出<strong>直观解释</strong>:Fisher 信息是<strong>对数似然在真值附近的曲率</strong>——曲率越大(似然越尖锐),参数越容易被数据确定,信息量越大。<strong>可加性</strong>:对独立样本,I_n(θ)=n·I(θ)(信息相加),这直接导致标准误 ∝1/√n。<strong>Cramér-Rao 下界(CRB)</strong>:对任何<strong>无偏</strong>估计量,Var(θ̂)≥1/(nI(θ))。它是无偏估计方差的<strong>理论下界</strong>,达到它的估计量称为<strong>有效估计量</strong>(efficient)。MLE 在正则条件下<strong>渐近达到 CRB</strong>(渐近有效),这是 MLE 最优性的核心依据。

🏭 Production Trade-offs

实践应用:① <strong>标准误与置信区间</strong>——渐近地 Var(θ̂_MLE)≈1/(nI(θ̂)),故标准误可用<strong>观测信息矩阵</strong>(海森的负逆)估计:SE=√[(−H)⁻¹]ᵢᵢ;这是逻辑回归、GLM、混合模型输出标准误的标准做法(<code>statsmodels</code> 即此)。② <strong>实验设计的指导</strong>——Fisher 信息给出'哪种实验设计能最大化参数信息'的判据:D-最优设计最大化 det(I(θ))、A-最优最小化 tr(I⁻¹),这是<strong>最优实验设计</strong>的理论基础。③ <strong>与海森矩阵的关系</strong>——在 MLE 处,观测信息 −H(θ̂) 与 Fisher 信息 I(θ̂) 在大样本下趋于一致;故'海森 = 信息',这解释了为什么牛顿法(用海森)在统计上与 Fisher 得分法(IRLS)一致。④ <strong>与 KL 散度的关系</strong>——两个邻近分布的 KL 散度局部近似为 ½·ΔθᵀI(θ)Δθ,即 Fisher 信息是<strong>统计流形上的度量张量</strong>(Riemannian metric),这是信息几何的基础,也用于自然梯度(用 I⁻¹ 调整梯度方向,实现参数空间的'最速下降')。⑤ <strong>局限</strong>——CRB 只适用于<strong>无偏</strong>估计;有偏估计(如岭回归、收缩估计)可以突破 CRB(用偏差换方差)。
⚠️ Common Interview Pitfalls
  • ✕
    把 CRB 当作所有估计量的下界(只适用于无偏估计)
  • ✕
    忽略 Fisher 信息的可加性(导致对 √n 规律的困惑)
🎯 Interviewer Follow-ups
  • ?
    Fisher 信息与海森矩阵的关系?
  • ?
    哪些估计量能达到 CRB?
📚

Associated Knowledge Base Guides & Mindmaps

Explore the comprehensive technical article, exam cards, and global architecture tree.

← PreviousM1-073: Numerical Stability: 解释 FP32 主权重(master weights)与它在混合精度训练中的作用。📋Back to BankNext →M1-075: 估计理论 (MLE/MAP): 什么是收缩估计(shrinkage)与 James-Stein 现象?