M4-058M4: Sequences & TransformersKV Cache & Inference OptimizationsEasy
Mastery:
KV Cache & Inference Optimizations: 解释 prefill 与 decode 两阶段的差异。
📐 Mathematical Definition
⚡ Executive Summary
Core Concept: prefill 并行处理输入(compute-bound、算力密集);decode 逐 token 生成(memory-bound、读 KV);两者需不同优化。
📌 Key Takeaways
- •prefill:整段输入并行,大矩阵乘,算力受限
- •decode:每步 1 token,读整个 KV cache,带宽受限
- •TTFT 由 prefill 决定;TPOT 由 decode 决定
📐 Mathematical Derivations
数学机理:<strong>prefill(预填充)</strong>——处理用户输入的整段 prompt(长度 L),所有位置的 Q/K/V 可<strong>一次性并行计算</strong>(类似训练的前向),注意力是 L×L 的大矩阵运算,FFN 是 L×d×d 的大矩阵乘;故 prefill 是 <strong>compute-bound</strong>(受算力限制),且能高效利用张量核心。<strong>decode(解码)</strong>——逐 token 生成,每步只有 1 个 query,需读取<strong>整个 KV cache</strong>(长度 S)做注意力;每步的 FLOPs 很小但访存量 ∝S,故是 <strong>memory-bound</strong>(受带宽限制),且并行度低(batch×heads)。<strong>关键指标</strong>:(a) <strong>TTFT(Time To First Token)</strong>——首 token 延迟,主要由 prefill 决定(∝ prompt 长度);(b) <strong>TPOT(Time Per Output Token)</strong>——每输出 token 的延迟,由 decode 决定(∝ 读取 KV 的量);(c) <strong>吞吐</strong>——由两阶段的效率与批大小共同决定。<strong>为什么需不同优化</strong>——prefill 适合<strong>大 batch 的矩阵乘</strong>(算力密集,batch 越大越好);decode 适合<strong>大 batch 以摊薄 KV 读取</strong>(带宽密集,batch 越大 KV 复用越高)。故推理引擎需<strong>分别调优</strong>,甚至<strong>分离部署</strong>(PD 分离)。<strong>两阶段的显存需求也不同</strong>——prefill 需大量激活显存(大矩阵中间结果),decode 需大量 KV cache 显存。
🏭 Production Trade-offs
深度剖析与工程权衡:① <strong>'chunked prefill'</strong>——把长 prompt 的 prefill 切成多个 chunk 逐步处理,与 decode 请求<strong>混批</strong>,使 (a) 长 prompt 不阻塞短请求(改善延迟)、(b) 计算与带宽资源互补(prefill 算力密集、decode 带宽密集,混批可同时利用);这是现代推理引擎(vLLM、TensorRT-LLM)的关键调度技术。② <strong>PD 分离(disaggregation)</strong>——把 prefill 与 decode 部署在<strong>不同实例</strong>(各自用最适合的硬件/并行策略),通过高速网络传递 KV cache;优点是各阶段可独立扩缩容、避免相互干扰,代价是 KV 传输开销。③ <strong>batch 策略的差异</strong>——prefill 用'大 batch + 长序列'填满算力;decode 用'尽量大的 batch'摊薄 KV 读取;两者的最优 batch 大小不同,故混批需调度算法(如 Sarathi-Serve 的 chunked prefill + piggybacking)。④ <strong>与投机解码的关系</strong>——投机解码在 decode 阶段一次验证多个 token,把'每步 1 个 query'变成'每步 k 个 query',从而提升 decode 的算术强度(更接近 prefill 的效率);这是'把 decode 变 compute-bound'的思路。⑤ <strong>与 KV 压缩的关系</strong>——decode 的瓶颈是读 KV,故 MQA/GQA/MLA/KV 量化直接改善 TPOT。⑥ <strong>面试要点</strong>——被问'prefill 与 decode 的区别',应给出'<strong>并行度(L vs 1)+ 瓶颈(算力 vs 带宽)+ 指标(TTFT vs TPOT)+ 优化方向</strong>'四维对比,并说明'chunked prefill / PD 分离'的工程动机;这是推理系统类问题的核心考点。
⚠️ Common Interview Pitfalls
- ✕把 prefill 与 decode 的瓶颈当作同类
- ✕忽略 TTFT 与 TPOT 的区分
🎯 Interviewer Follow-ups
- ?为什么 prefill 与 decode 需要不同的 batch 策略?
- ?PD 分离的动机是什么?
📚
Associated Knowledge Base Guides & Mindmaps
Explore the comprehensive technical article, exam cards, and global architecture tree.