AI Roadmap/Layer 05 · 05. Training, Post-Training & Alignment
5.3

5.3 Preference Optimization, Alignment & Reasoning

DPO/IPO/SimPO offline preference optimization, Constitutional AI, PPO online RL, DeepSeek-R1 GRPO group relative sampling, and rule-based verifiable rewards (RLVR).

DPO / IPO Offline Preference Optimization & Data Curation

DPO closed-form optimization bypassing separate reward models, IPO/KTO/SimPO variants, Constitutional AI, preference dataset curation, and out-of-distribution (OOD) degeneration mitigation.

🏢 Companies
AnthropicOpenAIMeta (Llama)Google DeepMindScale AI
🛠️ Tech Stack
DPOIPOKTOSimPOConstitutional AIPreference PairsOOD Degeneration
💼 Roles & Salary
Alignment Engineer、Safety Researcher、RL Engineer
💰 $255K - $540K / year (Preference Alignment) | ¥800K - ¥2.0M / year
📚 Prerequisites: Bradley-Terry Preference Math Derivations • DPO Loss Function & Gradient Derivations • Reference Model Mechanics & Memory Tradeoffs • Human/AI Feedback Annotation Quality

PPO Online RL, DeepSeek GRPO & Rule-Verifiable Rewards (RLVR)

PPO online policy gradient with Actor-Critic/Reward models, DeepSeek-R1 GRPO group relative reward sampling (bypassing separate Critic networks), rule-based verifiable rewards (RLVR) for math/code reasoning, and Process Reward Models (PRM).

🏢 Companies
DeepSeekOpenAIAnthropicGoogle DeepMind
🛠️ Tech Stack
GRPODeepSeek-R1PPORLVRProcess Reward ModelsRule-Based RewardsActor-Critic
💼 Roles & Salary
RL Engineer、Alignment Engineer、Research Scientist、Evaluation Engineer
💰 $270K - $620K / year (Frontier RL & Reasoning) | ¥900K - ¥2.5M / year
📚 Prerequisites: Markov Decision Processes & Policy Gradients • DeepSeek-R1 GRPO Group Normalized Reward Math • PPO Clipped Objective & GAE Advantage Estimation • Code/Math Sandboxes for Verifiable Rewards (RLVR)