M1-044M1: Mathematics & Statistics FundamentalsHypothesis TestingMedium
Mastery:

Hypothesis Testing: 什么是多重比较问题?常见校正方法有哪些。

📐 Mathematical Definition
Bonferroni: α′=α/m,BH: p(i)≤imq\text{Bonferroni}:\ \alpha'=\alpha/m,\qquad \text{BH}:\ p_{(i)}\le\frac{i}{m}q
⚡ Executive Summary
Core Concept: 多次检验会抬高族错误率;校正方法控制 FWER 或 FDR。

📌 Key Takeaways

  • •
    Bonferroni 保守(控制 FWER)
  • •
    BH 控制 FDR,功效更高,适合大规模检验(基因组/特征筛选)

📐 Mathematical Derivations

问题的数学本质:m 个独立检验、每个用 α=0.05 时,至少一次假阳性的概率为 1−(1−α)^m;m=10 时约 40%,m=100 时约 99.4%。<strong>FWER(族错误率)</strong> = P(至少一次假阳性),<strong>FDR(错误发现率)</strong> = E[假阳性数/总发现数]。两者对应不同的控制目标:FWER 更严格(要求零假阳性),FDR 允许一定比例的假阳性但控制其期望。<strong>Bonferroni</strong> 校正把每个检验的阈值设为 α/m,由并集界 P(∪Aᵢ)≤ΣP(Aᵢ) 保证 FWER≤α——简单但<strong>极其保守</strong>(m 大时功效骤降)。<strong>Benjamini-Hochberg (BH)</strong> 把 p 值升序排列后,找最大的 i 使 p₍ᵢ₎≤(i/m)q,拒绝前 i 个假设——控制 FDR≤q,功效远高于 Bonferroni。

🏭 Production Trade-offs

选择依据:① <strong>FWER 适用场景</strong>——错误代价极高、发现数少(如临床试验的主要终点、安全相关指标);② <strong>FDR 适用场景</strong>——大规模筛选、允许一定假阳性(如基因组差异表达分析、特征筛选、异常检测)。在 A/B 测试中,多重比较的处理方式略有不同:通常<strong>指定唯一的 primary metric</strong>(避免 p-hacking),其余指标作为护栏/探索性指标不做显著性声称;若必须同时看多个指标,可用分层检验(gatekeeping)或把多个指标合成单一综合指标(如 OEC,Overall Evaluation Criterion)。此外,序贯检验(允许中途查看)也需专门的 alpha-spending 方法,不能简单套用 Bonferroni。
⚠️ Common Interview Pitfalls
  • ✕
    在 A/B 中对所有指标都做显著性声称(应用唯一主指标)
  • ✕
    在需要严格控制时使用 BH(它允许假阳性)
🎯 Interviewer Follow-ups
  • ?
    A/B 中同时看很多指标该怎么处理?
  • ?
    FWER 与 FDR 的区别?
📚

Associated Knowledge Base Guides & Mindmaps

Explore the comprehensive technical article, exam cards, and global architecture tree.

← PreviousM1-043: Hypothesis Testing: 样本量与功效的关系是什么?写出常用的样本量估算思路。📋Back to BankNext →M1-045: Hypothesis Testing: 什么是序贯检验与 peeking 问题?如何正确做序贯实验。