M4-063M4: Sequences & TransformersKV Cache & Inference OptimizationsHard
Mastery:

KV Cache & Inference Optimizations: 解释 chunked prefill 与混合批处理调度。

📐 Mathematical Definition
chunked:prefill split into c chunks;mix with decode⇒utilize both resources\text{chunked}: \text{prefill split into }c\ \text{chunks};\qquad \text{mix with decode}\Rightarrow\text{utilize both resources}
⚡ Executive Summary
Core Concept: 把长 prompt 的 prefill 切成小块,与 decode 请求混在同一批中,兼顾 TTFT 与吞吐(算力与带宽互补)。

📌 Key Takeaways

  • •
    长 prompt 分块处理,避免阻塞 decode 请求
  • •
    prefill(算力密集)与 decode(带宽密集)混批互补
  • •
    是 Sarathi-Serve / vLLM 等引擎的关键调度技术

📐 Mathematical Derivations

数学机理:<strong>问题</strong>——在连续批处理下,新请求的 <strong>prefill</strong>(处理长 prompt,长度可能数千 token)会<strong>独占算力</strong>,使正在进行的 decode 请求被阻塞、延迟抖动(decode 请求的 TPOT 变差)。同时,prefill 是 compute-bound(算力密集、访存少),而 decode 是 memory-bound(访存密集、算力空闲)——<strong>两者的资源需求互补</strong>,若分开执行则各自浪费一半资源。<strong>chunked prefill</strong> 的解法:把长 prompt 的 prefill <strong>切成固定大小的小块(chunk,如 512 token)</strong>,每次只处理一块,并把该块与<strong>同批中的 decode 请求一起执行</strong>。<strong>收益</strong>:(a) <strong>不阻塞 decode</strong>——每个 chunk 的计算量受控,decode 请求的延迟抖动降低;(b) <strong>资源互补</strong>——prefill 块提供算力密集的工作(填满张量核心)、decode 提供访存密集的工作(填满带宽),混批使 GPU 的两类资源都被利用;(c) <strong>TTFT 可控</strong>——chunk 大小决定'一个长 prompt 需要多少步完成 prefill',从而控制 TTFT 与吞吐的权衡。<strong>与 Sarathi-Serve 的关系</strong>——该工作(Agrawal 等 2024)系统化了'chunked prefill + piggybacking(把 prefill 块'搭载'在 decode 批上)',并证明其能同时改善吞吐与延迟抖动(称为'stall-free batching')。<strong>chunk 大小的权衡</strong>——chunk 越大则 prefill 越快完成(TTFT 低)但单步延迟抖动大;chunk 越小则抖动小但 TTFT 高(需更多步)。实践中按目标 SLO 选择。

🏭 Production Trade-offs

深度剖析与工程权衡:① <strong>'算力与带宽互补'是核心洞察</strong>——这是理解 chunked prefill 的钥匙;在 roofline 视角下,prefill 位于 compute-bound 区、decode 位于 memory-bound 区,混批使工作点更接近'屋顶'的拐点(两者资源都被利用)。② <strong>与 PD 分离的对比</strong>——chunked prefill 是'<strong>同实例混批</strong>'(简单、无 KV 传输开销);PD 分离是'<strong>异实例分离</strong>'(各阶段可独立优化与扩缩容,但有 KV 传输开销)。两者是同一目标(消除两阶段互相干扰)的两种实现,可结合(如 PD 分离 + 各自的 chunked 调度)。③ <strong>调度复杂度</strong>——需在每步决定'哪些 prefill 块与哪些 decode 请求组成批',且需保证 (a) 不超出显存(KV 与激活)、(b) 满足 TTFT/TPOT 的 SLO;这是推理引擎调度器的核心算法。④ <strong>与投机解码的交互</strong>——投机解码使 decode 每步的 token 数可变(变长),进一步增加批组装的复杂度。⑤ <strong>实测收益</strong>——Sarathi-Serve 报告在同等吞吐下可显著降低 TPOT 的 P99 抖动(从数十倍降到几倍),这对'在线服务'的用户体验很关键(平均延迟好但尾延迟差是常见问题)。⑥ <strong>面试要点</strong>——被问'如何同时优化 TTFT 与吞吐',应给出'<strong>chunked prefill + 与 decode 混批(算力/带宽互补)+ chunk 大小权衡</strong>',并说明'与 PD 分离是同一目标的两种实现';能提到'尾延迟(P99 TPOT)'是深度理解的标志。
⚠️ Common Interview Pitfalls
  • ✕
    把 prefill 与 decode 分开执行(浪费互补的算力/带宽)
  • ✕
    chunk 大小设得过小导致 TTFT 过高
🎯 Interviewer Follow-ups
  • ?
    为什么 prefill 与 decode 混批能互补?
  • ?
    chunk 大小如何影响 TTFT 与吞吐?
📚

Associated Knowledge Base Guides & Mindmaps

Explore the comprehensive technical article, exam cards, and global architecture tree.

← PreviousM4-062: KV Cache & Inference Optimizations: 如何降低 KV Cache 显存?列出主要方法。📋Back to BankNext →M4-064: KV Cache & Inference Optimizations: 解释投机解码的变体(Medusa / EAGLE / MTP)。