AI Roadmap/Layer 06 · 06. Inference Acceleration & Model Serving
6.2

6.2 KV Cache, Context & Decoding Optimization

PagedAttention virtual memory paging, Automatic Prefix Caching (APC), RadixAttention, Chunked Prefill, Disaggregated Prefill/Decode, and Speculative Decoding.

PagedAttention Virtual Paging & Automatic Prefix Caching

OS-style virtual memory paging for dynamic non-contiguous KV Cache allocation, mitigating external memory fragmentation, and zero-copy prefix caching (APC/RadixAttention).

🏢 Companies
vLLM CommunitySGLangAnyscaleNVIDIA (TensorRT-LLM)
🛠️ Tech Stack
PagedAttentionKV CachePrefix CachingRadixAttentionVirtual MemoryvLLMSGLang
💼 Roles & Salary
Inference Engineer、ML Systems Engineer、C++ Performance Engineer
💰 $230K - $500K / year (KV Cache & Memory Systems) | ¥680K - ¥1.6M / year
📚 Prerequisites: OS Virtual Memory Page Tables & TLB • KV Cache Analytical Memory Formula • Radix Tree Matching & LRU Cache Eviction • Modern C++ Memory Pools & Concurrency

Chunked Prefill, Disaggregated Serving & Speculative Decoding

Chunked prefill balancing TTFT and TPOT, Disaggregated Prefill-Decode serving (Mooncake architecture), Speculative Decoding (EAGLE/Medusa draft verification), and Attention Sinks streaming.

🏢 Companies
vLLM CommunitySGLangTogether AIAnyscaleGroq
🛠️ Tech Stack
Chunked PrefillDisaggregated ServingSpeculative DecodingMooncakeEAGLEAttention SinksTTFT / TPOT
💼 Roles & Salary
Inference Engineer、ML Systems Engineer、Distributed Serving Engineer
💰 $240K - $510K / year (Disaggregated & Speculative Inference) | ¥720K - ¥1.7M / year
📚 Prerequisites: Compute vs Bandwidth Bound Prefill/Decode • Speculative Decoding Acceptance Math • Cross-Node KV Cache Transport (RDMA) • StreamingLLM & Attention Sinks Boundaries