M4-066M4: Sequences & TransformersKV Cache & Inference OptimizationsHard
Mastery:
KV Cache & Inference Optimizations: 解释 PD 分离(prefill-decode disaggregation)。
📐 Mathematical Definition
⚡ Executive Summary
Core Concept: 把 prefill 与 decode 部署在不同实例(各自优化),通过高速网络传 KV cache;消除两阶段互相干扰。
📌 Key Takeaways
- •prefill 与 decode 的资源需求不同,混部署会互相干扰
- •分离后可独立扩缩容、各自用最优并行策略
- •代价:KV cache 的网络传输开销
📐 Mathematical Derivations
数学机理:<strong>动机</strong>——prefill(compute-bound,算力密集)与 decode(memory-bound,带宽密集)在同一实例上运行时会<strong>互相干扰</strong>:(a) prefill 的长 prompt 会占用大量算力,使 decode 的 TPOT 抖动(这是 chunked prefill 要缓解的);(b) 两者对'最优并行策略'的要求不同(prefill 适合大 batch + TP,decode 适合大 batch + 低延迟);(c) 两者对硬件的最优配置不同(prefill 需要高算力、decode 需要高带宽/大显存)。<strong>PD 分离(disaggregation,DistServe/Zhong 等 2024、Splitwise/Patel 等 2024)</strong> 把两阶段部署在<strong>不同实例</strong>:prefill 实例处理输入并生成 KV cache,通过高速网络(RDMA/IB)把 KV cache <strong>传输</strong>给 decode 实例;decode 实例只做自回归生成。<strong>收益</strong>:(a) <strong>消除干扰</strong>——各阶段的延迟互不影响(TPOT 抖动大幅降低);(b) <strong>独立扩缩容</strong>——按负载分别调整 prefill 与 decode 的实例数(如长 prompt 多则加 prefill);(c) <strong>各自最优配置</strong>——prefill 用高算力卡/大 TP、decode 用大显存/大 batch。<strong>代价</strong>——<strong>KV cache 传输</strong>:量级为 2×层数×n_kv×d_h×prompt_len×bytes(如 7B 模型 4k prompt 约 0.5~2 GB);需高速网络(RDMA)与传输优化(分块传输、与计算重叠、压缩)。<strong>何时收益最大</strong>——(a) 长 prompt(prefill 重)+ 长输出(decode 重)的混合负载;(b) 高负载、需精细 SLO 控制的在线服务;(c) 异构硬件可用时。<strong>低负载或短序列时</strong>,传输开销可能超过收益。
🏭 Production Trade-offs
深度剖析与工程权衡:① <strong>KV 传输的优化</strong>——(a) 用 RDMA/IB 而非 TCP(带宽与延迟);(b) <strong>分块流式传输</strong>(prefill 生成一块传一块,与 decode 的启动重叠,降低 TTFT);(c) <strong>KV 压缩后再传</strong>(量化/稀疏化,减少字节);(d) <strong>分层传输</strong>(只传 decode 需要的层,或按需传输)。② <strong>与 chunked prefill 的对比</strong>——chunked prefill 是'同实例混批'(无传输开销,但仍有资源竞争);PD 分离是'异实例'(无竞争,但有传输开销)。两者可组合(如 PD 分离 + 各自 chunked 调度)。③ <strong>与 MoE/EP 的交互</strong>——MoE 模型的 prefill 与 decode 对专家并行的需求不同,PD 分离使两者可独立优化;这是大 MoE 服务的常见架构。④ <strong>与'弹性伸缩'的关系</strong>——PD 分离使'按需扩缩容'成为可能(如白天 prefill 多、夜间 decode 多),提升集群利用率。⑤ <strong>实践现状</strong>——vLLM、SGLang、TensorRT-LLM 等已支持 PD 分离;但部署复杂度显著提高(需管理两个集群 + KV 传输链路),故主要用于大规模在线服务。⑥ <strong>面试要点</strong>——被问'PD 分离',应给出'<strong>两阶段资源需求不同 → 混部署互相干扰 → 分离后独立优化 + KV 传输代价</strong>'的权衡,并说明'长 prompt + 长输出的混合负载下收益最大';能提到'KV 分块流式传输降低 TTFT'是深度理解的标志。
⚠️ Common Interview Pitfalls
- ✕以为 PD 分离总是更快(低负载时传输开销可能超过收益)
- ✕忽略 KV 传输的带宽需求(需 RDMA)
🎯 Interviewer Follow-ups
- ?PD 分离在什么负载下收益最大?
- ?KV 传输量有多大?如何优化?
📚
Associated Knowledge Base Guides & Mindmaps
Explore the comprehensive technical article, exam cards, and global architecture tree.