Back to ML Engineer Mind Map
中文·English
💻 ML EngineerID: mle-learning-rate-schedule-warmup

LR Warmup & Cosine Annealing Tuning

学习率 Warmup 与退火调优排障
🎯Core Definition
Learning Rate Warmup & Cosine Annealing Dynamic Scheduling governs the temporal trajectory of optimization step sizes in deep neural networks and LLMs to ensure early stability and late-stage convergence into flat minima; it operates in two contiguous phases: 1) Linear Warmup: ramping the learning rate from 0 to ηmax\eta_{\max} over the initial NwarmupN_{\text{warmup}} steps, preventing chaotic updates when Adam's second-moment variance estimates vtv_t are still uncalibrated; 2) Cosine Annealing Decay: smoothly decaying the learning rate along a half-cosine period ηt=ηmin+12(ηmaxηmin)(1+cos(tNwarmupTtotalNwarmupπ))\eta_t = \eta_{\min} + \frac{1}{2}(\eta_{\max} - \eta_{\min})(1 + \cos(\frac{t - N_{\text{warmup}}}{T_{\text{total}} - N_{\text{warmup}}} \pi)) down to ηmin\eta_{\min}, guiding weights into wide, robust basins.
💡Use Cases
Large language model pre-training, fine-tuning Vision Transformers, and resolving early loss spikes or training stagnation.
Key Problems Solved
Starting immediately with high learning rates causes early gradient instability and NaN crashes; step decays cause sharp gradient shocks; Warmup and Cosine schedules yield seamless, optimal convergence.
🎯5 High-Frequency Exam Points
1
Derive the Cosine Annealing schedule equation and explain why smooth half-cosine decay favors flat minima over abrupt step decays?
2
Why is Linear Warmup mandatory with AdamW (explaining uncalibrated second-moment variance vtv_t in early iterations)?
3
Explain the theoretical justification and breakdown limits of the Linear Scaling Rule (scaling learning rate proportionally with batch size)?
4
Explain Leslie Smith's 1Cycle policy and how inverse momentum coupling achieves Super-Convergence in fewer epochs?
5
How should learning rates be rewarmed and decayed during Continual Pre-training on new domain corpora?
🔗Foundational Prerequisite Cards (Click to Review)
Updated 2026-08-14
🎯
Test Your Knowledge: Practice Questions for "LR Warmup & Cosine Annealing Tuning"
Single choice pitfall questions with instant feedback and mistake tracking.
🚀 Start Card Practice
Previous CardPR-AUC vs ROC & Threshold TuningNext CardDistributed DDP Bottlenecks & Tuning

🔗 More ML Engineer Knowledge Cards

Bias-Variance Tradeoff & OverfittingLoss Function Taxonomy & GradientsOptimizer Convergence & MomentumEnsemble Stacking & Blending