Learning Rate Warmup & Cosine Annealing Dynamic Scheduling governs the temporal trajectory of optimization step sizes in deep neural networks and LLMs to ensure early stability and late-stage convergence into flat minima; it operates in two contiguous phases: 1) Linear Warmup: ramping the learning rate from 0 to
ηmax over the initial
Nwarmup steps, preventing chaotic updates when Adam's second-moment variance estimates
vt are still uncalibrated; 2) Cosine Annealing Decay: smoothly decaying the learning rate along a half-cosine period
ηt=ηmin+21(ηmax−ηmin)(1+cos(Ttotal−Nwarmupt−Nwarmupπ)) down to
ηmin, guiding weights into wide, robust basins.