M3-041M3: Deep Learning FoundationsLearning Rate SchedulesEasy
Mastery:

Learning Rate Schedules: 比较 cosine、step、linear、WSD 调度。

📐 Mathematical Definition
cosine: ηt=ηmin⁡+12(η0−ηmin⁡)(1+cos⁡πtT)\text{cosine}:\ \eta_t=\eta_{\min}+\tfrac12(\eta_0-\eta_{\min})\left(1+\cos\tfrac{\pi t}{T}\right)
⚡ Executive Summary
Core Concept: cosine 平滑衰减到 0、step 分段下降、linear 线性衰减、WSD 先稳定后快速衰减;LLM 主流是 cosine 或 WSD。

📌 Key Takeaways

  • •
    cosine 平滑无突变、末端接近 0,实践效果稳定
  • •
    step 需手工设 milestone,对超参敏感
  • •
    WSD 的 stable 段可随时分支/续训,衰减段决定最终质量

📐 Mathematical Derivations

数学机理:<strong>step decay</strong> 在预设 epoch 处把 lr 乘以 γ(如每 30 epoch ×0.1);优点是简单,缺点是下降点与幅度需手工设定、且在下降点处 loss 会突变。<strong>linear decay</strong> 从 η₀ 线性降到 η_min,衰减平稳但末端下降过快、可能过早'冻结'优化。<strong>cosine</strong> 用余弦曲线从 η₀ 平滑降到 η_min:前期下降慢(保持探索)、中后期加速下降(精细收敛)、末端变化率趋于 0(平稳收敛);其'末端导数趋零'的性质使训练后期非常稳定,是 CV/NLP 的通用默认。<strong>WSD(Warmup-Stable-Decay)</strong> 把训练分成三段:warmup(升)、stable(恒定 η₀,占大部分步数)、decay(快速降到接近 0,占约 10%~20%)。其理论依据来自'学习率与损失的关系'研究(如 WSD 论文与 μP 系列):在 stable 段模型持续在'高 lr 的高损失平台'上探索,最终质量主要由<strong>最后一段快速衰减</strong>决定——衰减把参数从'探索状态'精炼到'收敛状态'。

🏭 Production Trade-offs

深度剖析与工程权衡:① <strong>WSD 的工程优势</strong>——stable 段可随时保存 checkpoint 并在需要时'接着衰减',无需预先确定总步数;这对算力不确定或需持续预训练的场景极有价值。cosine 必须预先知道总步数,中途改变需重算曲线。② <strong>多阶段续训</strong>——WSD 让'预训练 → 继续预训练 → 退火'成为自然流程:多个 stable 段拼接,最后统一 decay;LLaMA-3 等即采用类似思路(cosine 与 WSD 的混合)。③ <strong>末端 lr 的重要性</strong>——把最终 lr 降到 η₀ 的 0~10% 是'精炼'的关键;若末端 lr 仍高,模型会停在噪声较大的状态。这也是'why decay to zero'的答案。④ <strong>step decay 的现代地位</strong>——在 CV 中仍有使用(如每 30 epoch ×0.1),但被 cosine 大量取代;LLM 几乎不用 step。⑤ <strong>lr 与 batch 的耦合</strong>——调度应结合 batch size 一起设计:大 batch 需要更大 lr 与更充分的 warmup;小 batch 则噪声本身提供正则。⑥ <strong>面试要点</strong>——若被问'你选哪个',给出理由链:'需随时续训 → WSD;步数确定且求稳 → cosine;简单快速原型 → step/linear'。
⚠️ Common Interview Pitfalls
  • ✕
    认为 cosine 一定优于其他(取决于是否需中途续训)
  • ✕
    末端 lr 不降到位(最终质量受损)
🎯 Interviewer Follow-ups
  • ?
    为什么 LLM 常用 cosine 而不是 step?
  • ?
    WSD 的 decay 段为什么可以很短?
📚

Associated Knowledge Base Guides & Mindmaps

Explore the comprehensive technical article, exam cards, and global architecture tree.

← PreviousM3-040: Learning Rate Schedules: 为什么需要 warmup?📋Back to BankNext →M3-042: Learning Rate Schedules: 学习率与 batch size 的关系是什么?线性缩放规则。