M8-032M8: ML Systems, Engineering & ResearchInference Serving & DeploymentMedium
Mastery:
Inference Serving & Deployment: 解释推理引擎(vLLM/Triton/TensorRT)的作用。
📐 Mathematical Definition
⚡ Executive Summary
Core Concept: Triton 是多框架服务编排;TensorRT 是算子级图优化与量化;vLLM 专攻 LLM 的高吞吐(PagedAttention)。
📌 Key Takeaways
- •Triton:多模型/多框架的服务编排、动态批处理、并发模型执行
- •TensorRT:图优化(算子融合)+ 量化 + kernel 自动调优
- •vLLM:LLM 专用(PagedAttention、连续批处理、前缀缓存)
📐 Mathematical Derivations
数学机理:<strong>三种推理引擎的定位</strong>——(1) <strong>Triton Inference Server</strong>(NVIDIA)——(a) <strong>定位</strong>——<strong>多模型/多框架的服务编排</strong>(serving orchestration);(b) <strong>能力</strong>——(i) 支持多框架(TensorRT/PyTorch/ONNX/TensorFlow);(ii) <strong>动态批处理</strong>(把并发请求合成 batch);(iii) <strong>并发模型执行</strong>(同一 GPU 上多模型并行);(iv) <strong>模型编排</strong>(ensemble/pipeline——如'预处理 → 模型 → 后处理');(v) 指标与健康检查;(c) <strong>场景</strong>——多模型服务、需要编排的复杂流水线。(2) <strong>TensorRT</strong>(NVIDIA)——(a) <strong>定位</strong>——<strong>算子级图优化与量化</strong>;(b) <strong>能力</strong>——(i) <strong>图优化</strong>(算子融合、常量折叠、内存复用);(ii) <strong>kernel 自动调优</strong>(针对具体 GPU 与 shape 选择最优 kernel);(iii) <strong>量化</strong>(INT8/FP8 + 校准);(iv) <strong>动态 shape</strong> 支持;(c) <strong>效果</strong>——相比朴素 PyTorch 推理可提速 2~5 倍;(d) <strong>场景</strong>——追求极致延迟/吞吐的<strong>单模型</strong>部署(如 CNN/传统模型);(e) <strong>缺点</strong>——(i) 编译时间长(引擎构建);(ii) 与具体 GPU 绑定(需为每个 GPU 架构编译);(iii) LLM 支持不如 vLLM(虽 TensorRT-LLM 已补齐)。(3) <strong>vLLM</strong>——(a) <strong>定位</strong>——<strong>LLM 专用高吞吐引擎</strong>;(b) <strong>核心</strong>——(i) <strong>PagedAttention</strong>(KV cache 分页管理——消除碎片);(ii) <strong>连续批处理</strong>(continuous batching——动态组批);(iii) <strong>前缀缓存</strong>(prefix caching / RadixAttention);(iv) <strong>投机解码</strong>支持;(v) 量化支持;(c) <strong>效果</strong>——相比朴素实现吞吐提升数倍到数十倍;(d) <strong>场景</strong>——LLM 推理(尤其高并发服务)。(4) <strong>其他</strong>——(a) <strong>TensorRT-LLM</strong>(NVIDIA 的 LLM 引擎——融合 TensorRT 优化 + LLM 特性);(b) <strong>SGLang</strong>(RadixAttention + 结构化输出);(c) <strong>TGI</strong>(HuggingFace);(d) <strong>ONNX Runtime</strong>(跨平台);(e) <strong>llama.cpp</strong>(CPU/边缘)。<strong>三者的关系</strong>——(a) <strong>Triton 是'服务层'</strong>(编排、批处理、多模型);(b) <strong>TensorRT 是'编译优化层'</strong>(图优化、kernel);(c) <strong>vLLM 是'LLM 专用引擎'</strong>;(d) <strong>可组合</strong>(如 Triton + TensorRT-LLM、Triton + vLLM backend)。<strong>选择依据</strong>——(a) <strong>LLM 服务</strong> → vLLM/SGLang/TensorRT-LLM;(b) <strong>多模型/多框架</strong> → Triton(编排);(c) <strong>极致延迟的 CV 模型</strong> → TensorRT;(d) <strong>跨平台/边缘</strong> → ONNX Runtime/llama.cpp。<strong>与其他问题的关系</strong>——(a) 与 M4 的'推理优化'(PagedAttention/连续批处理);(b) 与'成本与延迟优化';(c) 与'自动扩缩容'。<strong>实践建议</strong>——(a) <strong>LLM 用 vLLM/SGLang</strong>(生态成熟);(b) <strong>多模型编排用 Triton</strong>;(c) <strong>CV/极致延迟用 TensorRT</strong>;(d) <strong>可组合</strong>(Triton + 后端引擎);(e) <strong>注意'编译时间与 GPU 绑定'</strong>(TensorRT 的代价)。<strong>度量</strong>——(a) 延迟(TTFT/TPOT);(b) 吞吐;(c) 显存效率;(d) 编译/启动时间。
🏭 Production Trade-offs
深度剖析与工程权衡:① <strong>'Triton 是服务层、TensorRT 是编译层、vLLM 是 LLM 引擎'</strong>——三者的层次不同;面试中能区分是深度理解的标志。② <strong>'可组合'</strong>——Triton + vLLM backend 是常见组合。③ <strong>'TensorRT 的代价'</strong>——编译时间长 + GPU 绑定(需为每个架构编译)。④ <strong>'vLLM 的核心是 PagedAttention + 连续批处理'</strong>——这两个是吞吐提升的关键。⑤ <strong>'LLM 与 CV 的引擎选择不同'</strong>——LLM 用 vLLM(KV cache 管理),CV 用 TensorRT(图优化)。⑥ <strong>面试要点</strong>——被问'推理引擎怎么选',应给出'<strong>三者定位(编排/编译/LLM 引擎)+ 组合使用 + 选择依据(LLM vs CV vs 多模型)+ TensorRT 的代价</strong>';能指出'三者层次不同'是深度理解的标志。
⚠️ Common Interview Pitfalls
- ✕用 TensorRT 部署 LLM(不如 vLLM 成熟)
- ✕忽略 TensorRT 的编译时间与 GPU 绑定代价
🎯 Interviewer Follow-ups
- ?三者的定位差异?
- ?什么场景用 TensorRT?
📚
Associated Knowledge Base Guides & Mindmaps
Explore the comprehensive technical article, exam cards, and global architecture tree.