Back to AI Roadmap
LAYER 06

06. Inference Acceleration & Model Serving

Benchmark Companies:
TensorRT-LLMTensorRT-LLMvLLMvLLMSGLangSGLangOpenAI TritonOpenAI TritonAMDAMDIntelIntelModularModularAnyscaleAnyscaleGroqGroqCerebrasCerebrasFireworks AIFireworks AITogether AITogether AIModalModalTGITGIllama.cppllama.cppMIT HAN LabMIT HAN LabNeural MagicNeural MagicCentMLCentMLNVIDIANVIDIAMeta (Llama)Meta (Llama)ModularModularMIT HAN LabMIT HAN LabHugging FaceHugging FaceAMDAMDBasetenBaseten

Turning laboratory checkpoints into ultra-low-latency, memory-efficient, high-throughput, and resilient production serving engines.

📊Layer View Mode:
🌐3D CROSS-CORRELATION MATRIX

06. Inference Acceleration & Model Serving · Cross-Correlation Ecosystem

Click any company, tech category, or career role to illuminate all cross-associations and dim unrelated entities.

🔍
🏢

Benchmark Enterprises

25
TensorRT-LLMTensorRT-LLM
6
vLLMvLLM
3
SGLangSGLang
3
OpenAI TritonOpenAI Triton
1
AMDAMD
2
IntelIntel
2
ModularModular
2
AnyscaleAnyscale
3
GroqGroq
3
CerebrasCerebras
0
Fireworks AIFireworks AI
1
Together AITogether AI
2
ModalModal
1
TGITGI
2
llama.cppllama.cpp
1
MIT HAN LabMIT HAN Lab
2
Neural MagicNeural Magic
1
CentMLCentML
2
NVIDIANVIDIA
6
Meta (Llama)Meta (Llama)
1
ModularModular
2
MIT HAN LabMIT HAN Lab
2
Hugging FaceHugging Face
2
AMDAMD
2
BasetenBaseten
1
🛠️

Tech Stack Categories & Atomic Nodes

4 Tracks
6.1

6.1 Kernel, Compiler & Hardware Acceleration

CUDA C/C++ kernels, OpenAI Triton language, FlashAttention-1/2/3 IO-aware tiling, TensorRT graph compiler, AMD ROCm, and Intel OpenVINO heterogeneous acceleration.
Sub-Page
FlashAttention-1/2/3 IO Optimization & OpenAI Triton Kernels
💰 $250K - $540K / year (CUDA & Kernel Engineering) | ¥750K - ¥1.8M / year
IO-aware tiling, on-chip SRAM online softmax calculation, reducing HBM IO from $O(N^2)$ to $O(N)$ with standard $O(N^2 d)$ FLOPs, warp specialization, and OpenAI Triton kernel programming.
Benchmark:NVIDIANVIDIAOpenAI TritonOpenAI TritonMeta (Llama)Meta (Llama)ModularModularCentMLCentML
FlashAttention-3CUDAOpenAI TritonOnline SoftmaxTilingSRAMWarp SpecializationOperator Fusion
Deep Learning Compilers, Graph Fusion & Heterogeneous Hardware
💰 $230K - $490K / year (Compiler & Heterogeneous Acceleration) | ¥700K - ¥1.65M / year
TensorRT/TVM/XLA computational graph optimization, automatic operator fusion, CUTLASS 3.x GEMM templates, AMD ROCm/HIP, Intel OpenVINO, and Roofline bottleneck analysis.
Benchmark:NVIDIANVIDIAAMDAMDIntelIntelModularModularCentMLCentML
TensorRTTVMXLACUTLASSROCmOpenVINOOperator FusionRoofline Model
6.2

6.2 KV Cache, Context & Decoding Optimization

PagedAttention virtual memory paging, Automatic Prefix Caching (APC), RadixAttention, Chunked Prefill, Disaggregated Prefill/Decode, and Speculative Decoding.
Sub-Page
PagedAttention Virtual Paging & Automatic Prefix Caching
💰 $230K - $500K / year (KV Cache & Memory Systems) | ¥680K - ¥1.6M / year
OS-style virtual memory paging for dynamic non-contiguous KV Cache allocation, mitigating external memory fragmentation, and zero-copy prefix caching (APC/RadixAttention).
Benchmark:vLLMvLLMSGLangSGLangAnyscaleAnyscaleTensorRT-LLMTensorRT-LLM
PagedAttentionKV CachePrefix CachingRadixAttentionVirtual MemoryvLLMSGLang
Chunked Prefill, Disaggregated Serving & Speculative Decoding
💰 $240K - $510K / year (Disaggregated & Speculative Inference) | ¥720K - ¥1.7M / year
Chunked prefill balancing TTFT and TPOT, Disaggregated Prefill-Decode serving (Mooncake architecture), Speculative Decoding (EAGLE/Medusa draft verification), and Attention Sinks streaming.
Benchmark:vLLMvLLMSGLangSGLangTogether AITogether AIAnyscaleAnyscaleGroqGroq
Chunked PrefillDisaggregated ServingSpeculative DecodingMooncakeEAGLEAttention SinksTTFT / TPOT
6.3

6.3 Model Compression & Quantization

Hopper/Blackwell native FP8/NVFP4 compute, SmoothQuant/QuaRot outlier mitigation, AWQ activation-aware quantization, KV Cache compression, and quality gates.
Sub-Page
Native FP8 / NVFP4 Compute, SmoothQuant & GPTQ
💰 $210K - $440K / year (Model Quantization) | ¥600K - ¥1.4M / year
Ada/Hopper native FP8 and Blackwell NVFP4 fine-grained block quantization, SmoothQuant outlier channel scaling, GPTQ inverse Hessian error compensation, and GGUF cross-platform quantization.
Benchmark:NVIDIANVIDIAMIT HAN LabMIT HAN LabNeural MagicNeural Magicllama.cppllama.cppHugging FaceHugging Face
FP8 (E4M3/E5M2)NVFP4SmoothQuantGPTQGGUFINT4Tensor Cores
AWQ Salient Activation, KV Cache Quantization & Quality Gates
💰 $205K - $430K / year (KV Cache Quant & Quality Gates) | ¥580K - ¥1.35M / year
AWQ protecting salient top-1% activation channels, FP8/INT8 KV Cache quantization for 2x concurrency, calibration dataset standards, automated multi-benchmark regression gates, and fallback rollback strategies.
Benchmark:NVIDIANVIDIAMIT HAN LabMIT HAN LabAMDAMDIntelIntel
AWQKV Cache QuantizationCalibration DatasetPerplexity RegressionQuality GatesMMLUAccuracy Fallback
6.4

6.4 Production Serving Engines, Control Plane & Reliability

vLLM, SGLang, and TensorRT-LLM engines, continuous batching, Multi-LoRA concurrent serving, TTFT/TPOT latency SLO governance, admission control, and canary rollbacks.
Sub-Page
High-Throughput Engines & Continuous Batching
💰 $235K - $510K / year (Inference Engine Architecture) | ¥700K - ¥1.65M / year
vLLM, SGLang, TensorRT-LLM, TGI architectures, iteration-level continuous batching, dynamic multi-LoRA concurrent mounting, distributed tensor parallel serving, and high-throughput SSE/gRPC streaming.
Benchmark:vLLMvLLMSGLangSGLangTensorRT-LLMTensorRT-LLMTGITGIGroqGroq
vLLMSGLangTensorRT-LLMContinuous BatchingMulti-LoRATGISSE / gRPCModel Serving
Serving Control Plane, SLO Governance & Elastic Reliability
💰 $225K - $480K / year (Serving Reliability & Control Plane) | ¥650K - ¥1.55M / year
TTFT and TPOT P95/P99 millisecond latency SLO observability, adaptive admission control and load shedding, multi-model gateway routing, GPU autoscaling, canary deployments, and automated zero-downtime rollback.
Benchmark:Fireworks AIFireworks AITogether AITogether AIAnyscaleAnyscaleModalModalBasetenBasetenGroqGroq
SLO GovernanceTTFT / TPOTAdmission ControlLoad SheddingAutoscalingCanary DeploymentRollbackCost per Token
💼

Career Track Roles

15
💼CUDA Engineer
1
💼GPU Kernel Engineer
1
💼Compiler Engineer
1
💼Hardware Acceleration Engineer
1
💼Inference Engineer
6
💼Quantization Engineer
2
💼Serving Reliability Engineer
1
💼C++ Performance Engineer
2
💼AI Platform Engineer
2
💼ML Systems Engineer
3
💼Distributed Serving Engineer
1
💼Model Optimization Engineer
1
💼Evaluation Engineer
1
💼Backend Performance Engineer
1
💼MLOps / SRE Architect
1
🔄Sub-Domain Sequential Path (4 stages):
6.1

6.1 Kernel, Compiler & Hardware Acceleration

Open Sub-Page

CUDA C/C++ kernels, OpenAI Triton language, FlashAttention-1/2/3 IO-aware tiling, TensorRT graph compiler, AMD ROCm, and Intel OpenVINO heterogeneous acceleration.

FlashAttention-1/2/3 IO Optimization & OpenAI Triton Kernels

IO-aware tiling, on-chip SRAM online softmax calculation, reducing HBM IO from $O(N^2)$ to $O(N)$ with standard $O(N^2 d)$ FLOPs, warp specialization, and OpenAI Triton kernel programming.

🏢 Companies
NVIDIANVIDIAOpenAI TritonOpenAI TritonMeta (Llama)Meta (Llama)ModularModularCentMLCentML
🛠️ Tech Stack
FlashAttention-3CUDAOpenAI TritonOnline SoftmaxTilingSRAMWarp SpecializationOperator Fusion
💼 Roles & Salary
CUDA Engineer、GPU Kernel Engineer、Inference Engineer、C++ Performance Engineer
💰 $250K - $540K / year (CUDA & Kernel Engineering) | ¥750K - ¥1.8M / year
【Mini-Sandbox】FlashAttention Memory Savings
Sequence Length:4,096 tokens
Naive Attention
1,024 MB
O(N²) 显存暴涨
FlashAttention
32 MB
节省 97% 显存 (O(N))

Deep Learning Compilers, Graph Fusion & Heterogeneous Hardware

TensorRT/TVM/XLA computational graph optimization, automatic operator fusion, CUTLASS 3.x GEMM templates, AMD ROCm/HIP, Intel OpenVINO, and Roofline bottleneck analysis.

🏢 Companies
NVIDIANVIDIAAMDAMDIntelIntelModularModularCentMLCentML
🛠️ Tech Stack
TensorRTTVMXLACUTLASSROCmOpenVINOOperator FusionRoofline Model
💼 Roles & Salary
Compiler Engineer、Hardware Acceleration Engineer、ML Systems Engineer
💰 $230K - $490K / year (Compiler & Heterogeneous Acceleration) | ¥700K - ¥1.65M / year
6.2

6.2 KV Cache, Context & Decoding Optimization

Open Sub-Page

PagedAttention virtual memory paging, Automatic Prefix Caching (APC), RadixAttention, Chunked Prefill, Disaggregated Prefill/Decode, and Speculative Decoding.

PagedAttention Virtual Paging & Automatic Prefix Caching

OS-style virtual memory paging for dynamic non-contiguous KV Cache allocation, mitigating external memory fragmentation, and zero-copy prefix caching (APC/RadixAttention).

🏢 Companies
vLLMvLLMSGLangSGLangAnyscaleAnyscaleTensorRT-LLMTensorRT-LLM
🛠️ Tech Stack
PagedAttentionKV CachePrefix CachingRadixAttentionVirtual MemoryvLLMSGLang
💼 Roles & Salary
Inference Engineer、ML Systems Engineer、C++ Performance Engineer
💰 $230K - $500K / year (KV Cache & Memory Systems) | ¥680K - ¥1.6M / year

Chunked Prefill, Disaggregated Serving & Speculative Decoding

Chunked prefill balancing TTFT and TPOT, Disaggregated Prefill-Decode serving (Mooncake architecture), Speculative Decoding (EAGLE/Medusa draft verification), and Attention Sinks streaming.

🏢 Companies
vLLMvLLMSGLangSGLangTogether AITogether AIAnyscaleAnyscaleGroqGroq
🛠️ Tech Stack
Chunked PrefillDisaggregated ServingSpeculative DecodingMooncakeEAGLEAttention SinksTTFT / TPOT
💼 Roles & Salary
Inference Engineer、ML Systems Engineer、Distributed Serving Engineer
💰 $240K - $510K / year (Disaggregated & Speculative Inference) | ¥720K - ¥1.7M / year
6.3

6.3 Model Compression & Quantization

Open Sub-Page

Hopper/Blackwell native FP8/NVFP4 compute, SmoothQuant/QuaRot outlier mitigation, AWQ activation-aware quantization, KV Cache compression, and quality gates.

Native FP8 / NVFP4 Compute, SmoothQuant & GPTQ

Ada/Hopper native FP8 and Blackwell NVFP4 fine-grained block quantization, SmoothQuant outlier channel scaling, GPTQ inverse Hessian error compensation, and GGUF cross-platform quantization.

🏢 Companies
NVIDIANVIDIAMIT HAN LabMIT HAN LabNeural MagicNeural Magicllama.cppllama.cppHugging FaceHugging Face
🛠️ Tech Stack
FP8 (E4M3/E5M2)NVFP4SmoothQuantGPTQGGUFINT4Tensor Cores
💼 Roles & Salary
Quantization Engineer、Model Optimization Engineer、Inference Engineer
💰 $210K - $440K / year (Model Quantization) | ¥600K - ¥1.4M / year

AWQ Salient Activation, KV Cache Quantization & Quality Gates

AWQ protecting salient top-1% activation channels, FP8/INT8 KV Cache quantization for 2x concurrency, calibration dataset standards, automated multi-benchmark regression gates, and fallback rollback strategies.

🏢 Companies
NVIDIANVIDIAMIT HAN LabMIT HAN LabAMDAMDIntelIntel
🛠️ Tech Stack
AWQKV Cache QuantizationCalibration DatasetPerplexity RegressionQuality GatesMMLUAccuracy Fallback
💼 Roles & Salary
Quantization Engineer、Evaluation Engineer、Inference Engineer
💰 $205K - $430K / year (KV Cache Quant & Quality Gates) | ¥580K - ¥1.35M / year
6.4

6.4 Production Serving Engines, Control Plane & Reliability

Open Sub-Page

vLLM, SGLang, and TensorRT-LLM engines, continuous batching, Multi-LoRA concurrent serving, TTFT/TPOT latency SLO governance, admission control, and canary rollbacks.

High-Throughput Engines & Continuous Batching

vLLM, SGLang, TensorRT-LLM, TGI architectures, iteration-level continuous batching, dynamic multi-LoRA concurrent mounting, distributed tensor parallel serving, and high-throughput SSE/gRPC streaming.

🏢 Companies
vLLMvLLMSGLangSGLangTensorRT-LLMTensorRT-LLMTGITGIGroqGroq
🛠️ Tech Stack
vLLMSGLangTensorRT-LLMContinuous BatchingMulti-LoRATGISSE / gRPCModel Serving
💼 Roles & Salary
Inference Engineer、AI Platform Engineer、Backend Performance Engineer
💰 $235K - $510K / year (Inference Engine Architecture) | ¥700K - ¥1.65M / year

Serving Control Plane, SLO Governance & Elastic Reliability

TTFT and TPOT P95/P99 millisecond latency SLO observability, adaptive admission control and load shedding, multi-model gateway routing, GPU autoscaling, canary deployments, and automated zero-downtime rollback.

🏢 Companies
Fireworks AIFireworks AITogether AITogether AIAnyscaleAnyscaleModalModalBasetenBasetenGroqGroq
🛠️ Tech Stack
SLO GovernanceTTFT / TPOTAdmission ControlLoad SheddingAutoscalingCanary DeploymentRollbackCost per Token
💼 Roles & Salary
Serving Reliability Engineer、AI Platform Engineer、MLOps / SRE Architect
💰 $225K - $480K / year (Serving Reliability & Control Plane) | ¥650K - ¥1.55M / year