M1-049M1: Mathematics & Statistics FundamentalsConfidence Intervals & BootstrapMedium
Mastery:

Confidence Intervals & Bootstrap: bootstrap 与置换检验(permutation test)分别适用于什么场景。

📐 Mathematical Definition
permutation:shuffle labels⇒null distribution\text{permutation}: \text{shuffle labels} \Rightarrow \text{null distribution}
⚡ Executive Summary
Core Concept: bootstrap 估计统计量的抽样分布(可给 CI);置换检验在 H0 下构造零分布(给 p 值)。

📌 Key Takeaways

  • •
    置换检验要求可交换性(H0 下分组无差异)
  • •
    两者都无需正态假设

📐 Mathematical Derivations

两者的<strong>目的与假设不同</strong>:<strong>Bootstrap</strong> 通过有放回重采样估计统计量的<strong>抽样分布</strong>,用途是构造置信区间、估计标准误;它不对 H₀ 作假设,模拟的是'从总体再抽样'。<strong>Permutation test</strong> 通过随机打乱分组标签(保持各组样本量与数据值不变)构造 <strong>H₀ 下的零分布</strong>,用途是计算 p 值;其有效性依赖<strong>可交换性</strong>(exchangeability)——即 H₀ 为真时,样本的分组标签是可交换的(任意排列同样可能)。两者的共同优势是<strong>无需分布假设</strong>(不依赖正态性),适合小样本与非常规统计量。

🏭 Production Trade-offs

选择与陷阱:① <strong>目的驱动选择</strong>——要区间用 bootstrap,要 p 值用 permutation;两者可以互补(先 permutation 得 p 值,再 bootstrap 得效应量 CI)。② <strong>置换检验的适用边界</strong>——要求处理分配与样本独立(如随机化实验);对<strong>观察性数据</strong>,若两组在混杂变量上不可交换,置换检验的零分布不再有效,此时需条件置换(在协变量层内置换)或用倾向性加权。③ <strong>依赖结构的破坏</strong>——时间序列(自相关)与聚类数据(用户内相关)打乱标签会破坏依赖结构,需用 block permutation 或 cluster permutation。④ <strong>A/B 测试中的重尾指标</strong>——两种方法都可用,但 permutation 对极值更稳健(因为它不依赖重采样复现尾部),实践中常先做 CUPED 降方差再检验;若指标是比率型(CTR),permutation 直接对用户级聚合值置换即可,无需正态假设。
⚠️ Common Interview Pitfalls
  • ✕
    对观察性数据直接用置换检验(违反可交换性)
  • ✕
    混淆 bootstrap(估分布)与 permutation(构零分布)的用途
🎯 Interviewer Follow-ups
  • ?
    A/B 中重尾指标该用哪个?
  • ?
    为什么置换检验不适用于复杂依赖结构?
📚

Associated Knowledge Base Guides & Mindmaps

Explore the comprehensive technical article, exam cards, and global architecture tree.

← PreviousM1-048: Confidence Intervals & Bootstrap: 什么是 bootstrap 的三种区间构造方法?📋Back to BankNext →M1-050: Confidence Intervals & Bootstrap: 如何用 bootstrap 做 A/B 测试的显著性判断?有什么坑。