AI Roadmap/Layer 04 · 04. Foundation Models & Model Assets
4.1

4.1 Foundation LLM & Reasoning Architectures

Transformer architecture evolutions, RoPE long-context scaling, Mixture-of-Experts (MoE) sparse routing, DeepSeek MLA attention, and chain-of-thought reasoning models.

Transformer Evolution & Long-Context Scaling

Dense Transformer foundations, KV cache architectures (MHA/GQA/MQA), RoPE position embeddings with YaRN/ALiBi long-context extrapolation, RMSNorm, SwiGLU, and BPE tokenizer designs.

🏢 Companies
OpenAIAnthropicGoogle DeepMindDeepSeekMoonshot AI (Kimi)智谱 AI (GLM)Qwen (Alibaba)Meta (Llama)Mistral AI
🛠️ Tech Stack
TransformerGQA / MQARoPEYaRNSwiGLURMSNormTokenizerLong-Context
💼 Roles & Salary
Research Scientist、ML Research Engineer、LLM Architect、NLP Engineer
💰 $260K - $560K / year (Frontier LLM Research) | ¥800K - ¥2.1M / year
📚 Prerequisites: PyTorch Internals & Transformer Math • Self-Attention FLOPs & Memory Complexity • KV Cache & GQA/MQA Matrix Derivations • Positional Encoding Theory (RoPE/YaRN)

MoE Sparse Routing & MLA Reasoning Architecture

Top-k sparse expert routing, auxiliary-loss-free dynamic load balancing, DeepSeek Multi-Head Latent Attention (MLA) low-rank KV compression, and chain-of-thought reasoning models.

🏢 Companies
DeepSeek智谱 AI (GLM)Moonshot AI (Kimi)Mistral AIMeta (Llama)Google DeepMindOpenAIQwen (Alibaba)
🛠️ Tech Stack
MoEExpert RoutingDeepSeek MLALow-Rank KVAuxiliary LossReasoning ModelDeepSeek-V3
💼 Roles & Salary
Research Scientist、LLM Architect、ML Research Engineer
💰 $270K - $600K / year (MoE & Reasoning Architectures) | ¥900K - ¥2.4M / year
📚 Prerequisites: MoE Gating Algorithms & Expert Parallelism • DeepSeek MLA Low-Rank KV Math Derivations • DeepSeek-V3 Dynamic Load Balancing • FlashAttention IO & Memory Optimization
【Mini-Sandbox】FlashAttention Memory Savings
Sequence Length:4,096 tokens
Naive Attention
1,024 MB
O(N²) 显存暴涨
FlashAttention
32 MB
节省 97% 显存 (O(N))