M4-111M4: Sequences & TransformersQuantization & AccelerationHard
Mastery:

Quantization & Acceleration: 如何系统评估一次推理优化的收益?

📐 Mathematical Definition
eval: bottleneck→metrics (TTFT/TPOT/throughput/goodput)→quality→ablation\text{eval}:\ \text{bottleneck}\to\text{metrics (TTFT/TPOT/throughput/goodput)}\to\text{quality}\to\text{ablation}
⚡ Executive Summary
Core Concept: 先定位瓶颈(roofline + profiling),再测端到端指标(TTFT/TPOT/吞吐/goodput)与质量(任务指标),并做消融与稳定性验证。

📌 Key Takeaways

  • •
    先定位瓶颈(memory-bound vs compute-bound、profiling)
  • •
    指标:TTFT、TPOT、吞吐、goodput(SLO 内)、显存
  • •
    同时验证质量(任务指标)与稳定性(P99 尾延迟)

📐 Mathematical Derivations

数学机理:<strong>评估的四步框架</strong>。<strong>(1) 定位瓶颈(先诊断再优化)</strong>——用 (a) <strong>roofline 分析</strong>(算算术强度,判断 memory-bound 还是 compute-bound);(b) <strong>profiling</strong>(Nsight/Torch Profiler 看各算子耗时占比、kernel 利用率、HBM 带宽利用);(c) <strong>分解阶段</strong>(prefill vs decode 的耗时与瓶颈不同)。<strong>没有定位就优化是盲目的</strong>(如对 compute-bound 的算子做稀疏化可能无效)。<strong>(2) 指标定义(多维度)</strong>——(a) <strong>TTFT</strong>(首 token 延迟,∝prefill);(b) <strong>TPOT</strong>(每 token 延迟,∝decode);(c) <strong>吞吐</strong>(tokens/s 或 requests/s);(d) <strong>goodput</strong>(在 SLO 约束下达到的有效吞吐——这才是服务的真实目标);(e) <strong>显存占用</strong>(权重 + KV + 激活);(f) <strong>尾延迟</strong>(P95/P99,而非只看均值——平均值好但尾部差是常见问题)。<strong>(3) 质量验证(不可忽略)</strong>——任何优化(量化/稀疏/蒸馏)都可能损质量;需在同一评测集上对比 (a) 困惑度、(b) 下游任务指标(含<strong>长上下文检索</strong>等敏感任务)、(c) 与全精度基线的差距;且需注意'PPL 平稳但检索能力下降'的陷阱。<strong>(4) 消融与稳定性</strong>——(a) <strong>消融实验</strong>(逐项开启优化,看每项的边际收益——避免'组合收益小于单项之和'的意外);(b) <strong>稳定性</strong>(长跑测试、不同输入分布下的表现、尾延迟);(c) <strong>成本</strong>(每百万 token 的硬件成本——最终决策依据)。<strong>常见陷阱</strong>——(a) <strong>只看吞吐不看延迟</strong>(可能违反 SLO);(b) <strong>只看均值不看尾延迟</strong>;(c) <strong>只看 PPL 不看任务指标</strong>;(d) <strong>不区分 prefill/decode</strong>(两者瓶颈不同、优化手段不同);(e) <strong>忽略 batch 大小的影响</strong>(同一优化在不同 batch 下收益不同)。<strong>实践方法</strong>——(a) 建立<strong>基线</strong>(未优化的 TTFT/TPOT/吞吐/质量);(b) 逐项优化并<strong>测量增量</strong>;(c) 在<strong>目标负载分布</strong>下测试(而非合成负载);(d) 记录<strong>成本-性能帕累托曲线</strong>(而非单点)。

🏭 Production Trade-offs

深度剖析与工程权衡:① <strong>'先定位瓶颈'是第一原则</strong>——优化前必须知道'瓶颈在哪'(带宽/算力/并行度/显存);否则可能优化了非瓶颈(Amdahl 定律:优化非瓶颈的收益 ≤ 其占比)。② <strong>goodput 是服务的真实指标</strong>——'最大吞吐'可能违反延迟 SLO;故应在 SLO 内最大化吞吐。这是从'实验室指标'到'生产指标'的关键转变。③ <strong>质量验证的敏感任务</strong>——长上下文检索(NIAH/RULER)对量化与稀疏最敏感;故评测集必须包含它们(否则会得出'无损'的错误结论)。④ <strong>消融的必要性</strong>——多个优化叠加时可能出现'相互干扰'(如量化 + 稀疏的误差叠加);故需逐项测量,避免'组合后反而更差'。⑤ <strong>负载分布的真实性</strong>——优化效果高度依赖负载(请求长度分布、并发数、是否共享前缀);故必须用<strong>生产负载的回放</strong>测试,而非合成数据。⑥ <strong>面试要点</strong>——被问'如何评估推理优化',应给出'<strong>定位瓶颈(roofline + profiling)→ 多维指标(TTFT/TPOT/吞吐/goodput/尾延迟)→ 质量验证(含敏感任务)→ 消融与稳定性 → 成本</strong>'的完整框架,并列出五个常见陷阱;这是'工程素养'的高分回答。
⚠️ Common Interview Pitfalls
  • ✕
    不做 profiling 就优化(可能优化非瓶颈)
  • ✕
    只看吞吐不看延迟 SLO 与尾延迟
🎯 Interviewer Follow-ups
  • ?
    为什么必须先定位瓶颈?
  • ?
    什么是 goodput?
📚

Associated Knowledge Base Guides & Mindmaps

Explore the comprehensive technical article, exam cards, and global architecture tree.

← PreviousM4-110: Quantization & Acceleration: 解释批处理、内核融合与算子优化对推理吞吐的影响。📋Back to BankNext →M4-112: Quantization & Acceleration: 解释 GPTQ / AWQ / SmoothQuant 三种量化算法的差异。