Back to AI Roadmap
🎯
LAYER 05

05. Training, Post-Training & Alignment

Benchmark Companies:
NVIDIANVIDIADeepSpeedDeepSpeedMeta (Llama)Meta (Llama)Google DeepMindGoogle DeepMindOpenAIOpenAIAnthropicAnthropicDeepSeek 深度求索DeepSeek 深度求索xAIxAITogether AITogether AICentMLCentMLHugging Face (PEFT)Hugging Face (PEFT)UnslothUnslothPredibasePredibaseScale AIScale AIDatabricksDatabricksSnorkel AISnorkel AIMeta (Llama)Meta (Llama)Hugging FaceHugging FaceMeta (Llama)Meta (Llama)

Training raw base models into reasoning-dense, instruction-following, human-aligned, and verifiable agents.

📊Layer View Mode:
🌐3D CROSS-CORRELATION MATRIX

05. Training, Post-Training & Alignment · Cross-Correlation Ecosystem

Click any company, tech category, or career role to illuminate all cross-associations and dim unrelated entities.

🔍
🏢

Benchmark Enterprises

19
NVIDIANVIDIA
2
DeepSpeedDeepSpeed
1
Meta (Llama)Meta (Llama)
3
Google DeepMindGoogle DeepMind
3
OpenAIOpenAI
3
AnthropicAnthropic
2
DeepSeek 深度求索DeepSeek 深度求索
3
xAIxAI
1
Together AITogether AI
2
CentMLCentML
1
Hugging Face (PEFT)Hugging Face (PEFT)
2
UnslothUnsloth
1
PredibasePredibase
1
Scale AIScale AI
2
DatabricksDatabricks
2
Snorkel AISnorkel AI
1
Meta (Llama)Meta (Llama)
2
Hugging FaceHugging Face
2
Meta (Llama)Meta (Llama)
2
🛠️

Tech Stack Categories & Atomic Nodes

3 Tracks
5.1

5.1 Large-Scale Training Systems

3D Parallelism (TP/PP/DP), Context Parallelism (CP), PyTorch FSDP, DeepSpeed ZeRO memory sharding, long-horizon fault tolerance, and async checkpoint recovery.
Sub-Page
3D Parallelism (TP/PP/DP), FSDP & ZeRO Memory Sharding
💰 $245K - $510K / year (Distributed Training) | ¥750K - ¥1.85M / year
Megatron-LM TP/PP communication scheduling, Context/Sequence Parallelism (CP/SP) for long context, PyTorch FSDP, DeepSpeed ZeRO-1/2/3 parameter/optimizer sharding, and activation checkpointing memory trade-offs.
Benchmark:NVIDIANVIDIADeepSpeedDeepSpeedMeta (Llama)Meta (Llama)DeepSeek 深度求索DeepSeek 深度求索Google DeepMindGoogle DeepMindOpenAIOpenAI
Tensor ParallelismPipeline ParallelismZeRO-3FSDPContext ParallelismMegatron-LMActivation Checkpointing
Training Fault Tolerance, Checkpointing & Observability
💰 $235K - $480K / year (Training Reliability) | ¥700K - ¥1.75M / year
10K+ GPU async checkpoint persistence and rapid recovery, straggler and silent hang detection, FP8/BF16 mixed-precision numerical stability, automated loss spike remediation, and W&B telemetry.
Benchmark:NVIDIANVIDIADeepSeek 深度求索DeepSeek 深度求索Meta (Llama)Meta (Llama)Together AITogether AICentMLCentMLxAIxAI
Fault ToleranceAsync CheckpointingStraggler DetectionLoss Spike RemediationFP8 StabilityTraining TelemetryW&B
5.2

5.2 Supervised Fine-Tuning & Adaptation

LoRA/QLoRA/DoRA low-rank adaptation, loss masking, multi-task data mixtures, curriculum learning, and offline quality regression gates.
Sub-Page
LoRA, QLoRA, DoRA & High-Throughput Fine-Tuning
💰 $195K - $390K / year (SFT & PEFT Engineering) | ¥500K - ¥1.2M / year
LoRA low-rank decomposition math ($W = W_0 + \frac{\alpha}{r}BA$), QLoRA NormalFloat4 (NF4) quantized fine-tuning, DoRA magnitude-direction decoupling, loss masking, and Unsloth/PEFT fast kernels.
Benchmark:UnslothUnslothHugging Face (PEFT)Hugging Face (PEFT)PredibasePredibaseTogether AITogether AIDatabricksDatabricks
LoRAQLoRADoRASFTLoss MaskingPEFTUnsloth
Instruction Mixture, Curriculum Tuning & Regression Gates
💰 $185K - $370K / year (Instruction & Eval Gates) | ¥480K - ¥1.1M / year
Multi-domain instruction weighting, difficulty-graded curriculum fine-tuning, catastrophic forgetting mitigation, chat template format safety, and automated pre-deployment regression gates.
Benchmark:Hugging FaceHugging FaceScale AIScale AIDatabricksDatabricksSnorkel AISnorkel AI
Instruction MixtureCurriculum TuningCatastrophic ForgettingChat TemplateRegression GatesEvaluation Gate
5.3

5.3 Preference Optimization, Alignment & Reasoning

DPO/IPO/SimPO offline preference optimization, Constitutional AI, PPO online RL, DeepSeek-R1 GRPO group relative sampling, and rule-based verifiable rewards (RLVR).
Sub-Page
DPO / IPO Offline Preference Optimization & Data Curation
💰 $255K - $540K / year (Preference Alignment) | ¥800K - ¥2.0M / year
DPO closed-form optimization bypassing separate reward models, IPO/KTO/SimPO variants, Constitutional AI, preference dataset curation, and out-of-distribution (OOD) degeneration mitigation.
Benchmark:AnthropicAnthropicOpenAIOpenAIMeta (Llama)Meta (Llama)Google DeepMindGoogle DeepMindScale AIScale AI
DPOIPOKTOSimPOConstitutional AIPreference PairsOOD Degeneration
PPO Online RL, DeepSeek GRPO & Rule-Verifiable Rewards (RLVR)
💰 $270K - $620K / year (Frontier RL & Reasoning) | ¥900K - ¥2.5M / year
PPO online policy gradient with Actor-Critic/Reward models, DeepSeek-R1 GRPO group relative reward sampling (bypassing separate Critic networks), rule-based verifiable rewards (RLVR) for math/code reasoning, and Process Reward Models (PRM).
Benchmark:DeepSeek 深度求索DeepSeek 深度求索OpenAIOpenAIAnthropicAnthropicGoogle DeepMindGoogle DeepMind
GRPODeepSeek-R1PPORLVRProcess Reward ModelsRule-Based RewardsActor-Critic
💼

Career Track Roles

13
💼Training Engineer
2
💼Distributed Training Engineer
2
💼ML Systems Engineer
1
💼Fine-Tuning Engineer
2
💼Alignment Engineer
2
💼RL Engineer
2
💼Safety Researcher
1
💼Evaluation Engineer
2
💼Research Scientist
1
💼Research Engineer
1
💼Performance Engineer
1
💼ML Engineer
2
💼LLM Application Engineer
1
🔄Sub-Domain Sequential Path (3 stages):
5.1

5.1 Large-Scale Training Systems

Open Sub-Page

3D Parallelism (TP/PP/DP), Context Parallelism (CP), PyTorch FSDP, DeepSpeed ZeRO memory sharding, long-horizon fault tolerance, and async checkpoint recovery.

3D Parallelism (TP/PP/DP), FSDP & ZeRO Memory Sharding

Megatron-LM TP/PP communication scheduling, Context/Sequence Parallelism (CP/SP) for long context, PyTorch FSDP, DeepSpeed ZeRO-1/2/3 parameter/optimizer sharding, and activation checkpointing memory trade-offs.

🏢 Companies
NVIDIANVIDIADeepSpeedDeepSpeedMeta (Llama)Meta (Llama)DeepSeek 深度求索DeepSeek 深度求索Google DeepMindGoogle DeepMindOpenAIOpenAI
🛠️ Tech Stack
Tensor ParallelismPipeline ParallelismZeRO-3FSDPContext ParallelismMegatron-LMActivation Checkpointing
💼 Roles & Salary
Training Engineer、Distributed Training Engineer、ML Systems Engineer
💰 $245K - $510K / year (Distributed Training) | ¥750K - ¥1.85M / year

Training Fault Tolerance, Checkpointing & Observability

10K+ GPU async checkpoint persistence and rapid recovery, straggler and silent hang detection, FP8/BF16 mixed-precision numerical stability, automated loss spike remediation, and W&B telemetry.

🏢 Companies
NVIDIANVIDIADeepSeek 深度求索DeepSeek 深度求索Meta (Llama)Meta (Llama)Together AITogether AICentMLCentMLxAIxAI
🛠️ Tech Stack
Fault ToleranceAsync CheckpointingStraggler DetectionLoss Spike RemediationFP8 StabilityTraining TelemetryW&B
💼 Roles & Salary
Distributed Training Engineer、Performance Engineer、Research Engineer
💰 $235K - $480K / year (Training Reliability) | ¥700K - ¥1.75M / year
5.2

5.2 Supervised Fine-Tuning & Adaptation

Open Sub-Page

LoRA/QLoRA/DoRA low-rank adaptation, loss masking, multi-task data mixtures, curriculum learning, and offline quality regression gates.

LoRA, QLoRA, DoRA & High-Throughput Fine-Tuning

LoRA low-rank decomposition math ($W = W_0 + \frac{\alpha}{r}BA$), QLoRA NormalFloat4 (NF4) quantized fine-tuning, DoRA magnitude-direction decoupling, loss masking, and Unsloth/PEFT fast kernels.

🏢 Companies
UnslothUnslothHugging Face (PEFT)Hugging Face (PEFT)PredibasePredibaseTogether AITogether AIDatabricksDatabricks
🛠️ Tech Stack
LoRAQLoRADoRASFTLoss MaskingPEFTUnsloth
💼 Roles & Salary
Fine-Tuning Engineer、ML Engineer、LLM Application Engineer
💰 $195K - $390K / year (SFT & PEFT Engineering) | ¥500K - ¥1.2M / year

Instruction Mixture, Curriculum Tuning & Regression Gates

Multi-domain instruction weighting, difficulty-graded curriculum fine-tuning, catastrophic forgetting mitigation, chat template format safety, and automated pre-deployment regression gates.

🏢 Companies
Hugging FaceHugging FaceScale AIScale AIDatabricksDatabricksSnorkel AISnorkel AI
🛠️ Tech Stack
Instruction MixtureCurriculum TuningCatastrophic ForgettingChat TemplateRegression GatesEvaluation Gate
💼 Roles & Salary
Fine-Tuning Engineer、Evaluation Engineer、ML Engineer
💰 $185K - $370K / year (Instruction & Eval Gates) | ¥480K - ¥1.1M / year
5.3

5.3 Preference Optimization, Alignment & Reasoning

Open Sub-Page

DPO/IPO/SimPO offline preference optimization, Constitutional AI, PPO online RL, DeepSeek-R1 GRPO group relative sampling, and rule-based verifiable rewards (RLVR).

DPO / IPO Offline Preference Optimization & Data Curation

DPO closed-form optimization bypassing separate reward models, IPO/KTO/SimPO variants, Constitutional AI, preference dataset curation, and out-of-distribution (OOD) degeneration mitigation.

🏢 Companies
AnthropicAnthropicOpenAIOpenAIMeta (Llama)Meta (Llama)Google DeepMindGoogle DeepMindScale AIScale AI
🛠️ Tech Stack
DPOIPOKTOSimPOConstitutional AIPreference PairsOOD Degeneration
💼 Roles & Salary
Alignment Engineer、Safety Researcher、RL Engineer
💰 $255K - $540K / year (Preference Alignment) | ¥800K - ¥2.0M / year

PPO Online RL, DeepSeek GRPO & Rule-Verifiable Rewards (RLVR)

PPO online policy gradient with Actor-Critic/Reward models, DeepSeek-R1 GRPO group relative reward sampling (bypassing separate Critic networks), rule-based verifiable rewards (RLVR) for math/code reasoning, and Process Reward Models (PRM).

🏢 Companies
DeepSeek 深度求索DeepSeek 深度求索OpenAIOpenAIAnthropicAnthropicGoogle DeepMindGoogle DeepMind
🛠️ Tech Stack
GRPODeepSeek-R1PPORLVRProcess Reward ModelsRule-Based RewardsActor-Critic
💼 Roles & Salary
RL Engineer、Alignment Engineer、Research Scientist、Evaluation Engineer
💰 $270K - $620K / year (Frontier RL & Reasoning) | ¥900K - ¥2.5M / year