AI Roadmap/Layer 06 · 06. Inference Acceleration & Model Serving
6.1

6.1 Kernel, Compiler & Hardware Acceleration

CUDA C/C++ kernels, OpenAI Triton language, FlashAttention-1/2/3 IO-aware tiling, TensorRT graph compiler, AMD ROCm, and Intel OpenVINO heterogeneous acceleration.

FlashAttention-1/2/3 IO Optimization & OpenAI Triton Kernels

IO-aware tiling, on-chip SRAM online softmax calculation, reducing HBM IO from $O(N^2)$ to $O(N)$ with standard $O(N^2 d)$ FLOPs, warp specialization, and OpenAI Triton kernel programming.

🏢 Companies
NVIDIAOpenAI (Triton)MetaModularCentML
🛠️ Tech Stack
FlashAttention-3CUDAOpenAI TritonOnline SoftmaxTilingSRAMWarp SpecializationOperator Fusion
💼 Roles & Salary
CUDA Engineer、GPU Kernel Engineer、Inference Engineer、C++ Performance Engineer
💰 $250K - $540K / year (CUDA & Kernel Engineering) | ¥750K - ¥1.8M / year
📚 Prerequisites: GPU Memory Hierarchy (Registers/SRAM/HBM) • CUDA C++ & Warp Shuffle Intrinsics • FlashAttention Online Softmax Derivation • OpenAI Triton Language & JIT Internals
【Mini-Sandbox】FlashAttention Memory Savings
Sequence Length:4,096 tokens
Naive Attention
1,024 MB
O(N²) 显存暴涨
FlashAttention
32 MB
节省 97% 显存 (O(N))

Deep Learning Compilers, Graph Fusion & Heterogeneous Hardware

TensorRT/TVM/XLA computational graph optimization, automatic operator fusion, CUTLASS 3.x GEMM templates, AMD ROCm/HIP, Intel OpenVINO, and Roofline bottleneck analysis.

🏢 Companies
NVIDIAAMD (ROCm)Intel (OpenVINO)Modular (MAX/Mojo)CentML
🛠️ Tech Stack
TensorRTTVMXLACUTLASSROCmOpenVINOOperator FusionRoofline Model
💼 Roles & Salary
Compiler Engineer、Hardware Acceleration Engineer、ML Systems Engineer
💰 $230K - $490K / year (Compiler & Heterogeneous Acceleration) | ¥700K - ¥1.65M / year
📚 Prerequisites: Computation Graph IR & Pass Optimizations • CUTLASS Templates & Tensor Core Layouts • AMD ROCm/HIP Compatibility Layer • Roofline Arithmetic Intensity Modeling