M4-043M4: Sequences & TransformersAttention Variants (MHA / MQA / GQA)Medium
Mastery:
Attention Variants (MHA / MQA / GQA): 解释 cross-attention 与 self-attention 的差异与用途。
📐 Mathematical Definition
⚡ Executive Summary
Core Concept: self-attention 的 Q/K/V 同源(序列内交互);cross-attention 的 Q 来自解码器、K/V 来自编码器(跨序列对齐)。
📌 Key Takeaways
- •self:序列内 token 互相交互,长度相同
- •cross:解码器查询编码器,长度可不同
- •cross 是 Enc-Dec 架构的信息桥梁;Dec-only 模型无 cross
📐 Mathematical Derivations
数学机理:<strong>self-attention</strong> 的 Q、K、V 都来自<strong>同一个序列</strong> X(Q=XW^Q、K=XW^K、V=XW^V),实现<strong>序列内</strong>的位置间信息交换(token mixing);序列长度前后一致(输出长度=输入长度)。<strong>cross-attention</strong> 的 Q 来自<strong>目标序列</strong> Y(如解码器的隐状态),K/V 来自<strong>源序列</strong> X(如编码器的输出):Attn=softmax((YW^Q)(XW^K)ᵀ/√d)(XW^V)。它实现<strong>跨序列</strong>的信息对齐——解码器每步'查询'源序列的相关部分,这是 Enc-Dec 架构(翻译/摘要/ASR)的<strong>信息桥梁</strong>。<strong>关键差异</strong>:(1) <strong>长度不同</strong>——Q 长度=|Y|、K/V 长度=|X|,注意力矩阵为 |Y|×|X|(非方阵);(2) <strong>方向性</strong>——cross-attention 无因果 mask(源序列已完整可见,无需屏蔽未来);(3) <strong>KV cache 复用</strong>——K/V 来自编码器,对<strong>所有解码步相同</strong>,故可只计算一次并缓存(无需随解码增长);这是 Enc-Dec 推理的一个优势。<strong>在 Dec-only 模型中的替代</strong>——Dec-only 无 cross-attention,源信息通过'拼接进上下文'(self-attention 内部交互)实现;代价是源序列也要被自回归地'读'一遍(更长的序列、更多的 KV cache),且源序列的注意力是<strong>因果</strong>的(不能双向理解源)。这就是 Enc-Dec 在'源需双向理解'任务上的结构优势所在。<strong>多模态中的 cross-attention</strong>——VLM 中常用'文本 query 查询图像特征'(如 Flamingo 的 gated cross-attention),把图像作为 K/V 注入 LLM。
🏭 Production Trade-offs
深度剖析与工程权衡:① <strong>KV cache 的差异</strong>——self-attention 的 KV cache 随解码<strong>增长</strong>(每步新增一个 K/V);cross-attention 的 KV cache <strong>固定</strong>(源序列的 K/V 只算一次),故 Enc-Dec 在长源 + 短输出的任务(摘要、检索增强生成)上更省显存。② <strong>Enc-Dec 与 Dec-only 的计算对比</strong>——Dec-only 把源拼进上下文后,源部分也要参与 O(L²) 注意力且被因果读取;Enc-Dec 用编码器双向处理源(一次 O(L²))+ 解码器 cross-attention(O(|Y|·|X|)),在'长源'场景更高效。③ <strong>交叉注意力的可解释性</strong>——cross-attention 权重可直接视为'输出对输入的对齐'(翻译中近似对角),比 self-attention 更易解读。④ <strong>多模态融合的三种方式</strong>——(a) <strong>cross-attention 注入</strong>(Flamingo、早期的 VLM);(b) <strong>前缀拼接</strong>(把图像特征作为前缀 token 输入 LLM,如 LLaVA);(c) <strong>统一 tokenization</strong>(原生多模态,如 Chameleon、Qwen-VL 的部分设计)。三者在训练难度、参数量、灵活性上各有取舍。⑤ <strong>与检索增强的关系</strong>——RAG 的'检索结果注入'也可用 cross-attention(RETRO 用 chunked cross-attention 检索邻居)或前缀拼接(更常见);前者可处理更长的检索内容。⑥ <strong>面试要点</strong>——被问'cross vs self attention',应给出'<strong>Q/K/V 来源(同源 vs 跨序列)+ 长度(相同 vs 不同)+ mask(因果 vs 无)+ KV cache(增长 vs 固定)</strong>'四维对比,并说明'Dec-only 用前缀拼接替代 cross-attention 的代价';能提到 VLM 的三种融合方式是加分。
⚠️ Common Interview Pitfalls
- ✕以为 cross-attention 需要因果 mask
- ✕忽略 cross-attention 的 KV cache 可固定复用
🎯 Interviewer Follow-ups
- ?Dec-only 模型如何替代 cross-attention?
- ?cross-attention 的 KV cache 如何管理?
📚
Associated Knowledge Base Guides & Mindmaps
Explore the comprehensive technical article, exam cards, and global architecture tree.