M3-027M3: Deep Learning FoundationsNormalization TechniquesHard
Mastery:

Normalization Techniques: 解释 QK-Norm 与注意力中的归一化技术。

📐 Mathematical Definition
Q^=RMSNorm(Q),K^=RMSNorm(K)\hat Q=\mathrm{RMSNorm}(Q),\qquad \hat K=\mathrm{RMSNorm}(K)
⚡ Executive Summary
Core Concept: 对 attention 的 Q、K 做归一化,稳定 logits 尺度;现代 LLM(Gemma/Qwen)采用。

📌 Key Takeaways

  • •
    防止 attention logits 随维度/深度爆炸
  • •
    对低精度训练与长上下文尤其重要

📐 Mathematical Derivations

问题背景:attention 的 logits 为 QKᵀ/√d_k。在深层网络中,Q、K 的范数可能随训练增长(尤其缺少归一化或使用小 ε 时),导致 logits 的<strong>尺度爆炸</strong>——表现为 softmax 饱和(注意力集中在单个 token、梯度消失)、训练不稳、以及低精度(FP16/BF16)下的溢出。<strong>QK-Norm 的解法</strong>:在计算 attention 之前,对 Q 与 K 分别施加归一化(通常是 RMSNorm,也可是 LN),使其范数受控:Q̂=RMSNorm(Q)、K̂=RMSNorm(K),再算 Q̂K̂ᵀ/√d_k。<strong>效果</strong>:(a) attention logits 的尺度被限制在合理范围(因为 Q̂、K̂ 的范数 ≈√d,logits ≈√d·cos 有界);(b) softmax 不会饱和,梯度更健康;(c) 训练更稳、可支持更大学习率与更低精度。<strong>采用情况</strong>:Gemma、Qwen 等现代 LLM 在 attention 中加了 QK-Norm(Gemma 用 RMSNorm 作用于 Q、K);一些视觉 Transformer 也用(如某些 ViT 变体)。<strong>与 LN 的区别</strong>:QK-Norm 只作用于 attention 的 Q/K(局部),而 LN 作用于整个 block 的输入(全局);两者互补(QK-Norm 解决 attention 内部的尺度问题,LN 解决 block 间的尺度问题)。

🏭 Production Trade-offs

实践要点:① <strong>为什么长上下文更需要</strong>——长序列下 attention logits 的方差随序列长度增长(更多 token 参与竞争),尺度爆炸更严重;QK-Norm 能显著改善长上下文的训练稳定性。② <strong>与温度/缩放的关系</strong>——QK-Norm 本质上替代了'手动调 √d_k 或加温度'的需求;有些实现用可学习的温度(如 Qwen2 的 logit soft-capping)进一步限制 logits 范围。③ <strong>低精度训练的关键作用</strong>——FP16/BF16 下 logits 溢出会导致 NaN;QK-Norm 使 logits 保持在安全范围,是低精度大模型训练的稳定化手段之一。④ <strong>其他注意力稳定化技术</strong>——(a) <strong>logit soft-capping</strong>(对 logits 施加 tanh 软截断,Gemma 2 用);(b) <strong>注意力权重 dropout</strong>;(c) <strong>更大的 ε</strong>(RMSNorm 的 ε 放大以提升 FP16 稳定性)。⑤ <strong>实现注意</strong>——QK-Norm 的归一化需在<strong>每个 head</strong> 内做(而非跨 head),且应在 reshape 之后、attention 之前;参数(γ)通常初始化为 1。⑥ <strong>实践建议</strong>——训练长上下文或低精度大模型时,建议启用 QK-Norm;它对精度的负面影响很小(可忽略),但稳定性收益明显。
⚠️ Common Interview Pitfalls
  • ✕
    长上下文训练不做 attention 归一化(logits 爆炸)
  • ✕
    QK-Norm 跨 head 归一化(应逐 head)
🎯 Interviewer Follow-ups
  • ?
    为什么 attention logits 会爆炸?
  • ?
    QK-Norm 与 LN 的区别?
📚

Associated Knowledge Base Guides & Mindmaps

Explore the comprehensive technical article, exam cards, and global architecture tree.

← PreviousM3-026: Normalization Techniques: 解释 Weight Standardization 与它在无 BN 网络中的作用。📋Back to BankNext →M3-028: Normalization Techniques: 解释 DeepNorm 与它如何支持超深 Transformer。