M5-027M5: NLP & Large Language ModelsInstruction Tuning & Supervised Fine-TuningMedium
Mastery:

Instruction Tuning & Supervised Fine-Tuning: 解释 chat template 与多轮对话的构造。

📐 Mathematical Definition
tmpl: ⟨sys⟩…⟨user⟩…⟨assistant⟩… ;loss on assistant turns\text{tmpl}:\ \langle\text{sys}\rangle\dots\langle\text{user}\rangle\dots\langle\text{assistant}\rangle\dots;\qquad \text{loss on assistant turns}
⚡ Executive Summary
Core Concept: chat template 定义角色标记与格式;多轮数据需构造历史并决定损失落在哪些轮,格式不一致会损害指令遵循。

📌 Key Takeaways

  • •
    template 定义角色边界(system/user/assistant 的特殊 token)
  • •
    多轮需构造历史;损失通常算在所有 assistant 轮
  • •
    训练与推理必须用同一 template(否则性能崩塌)

📐 Mathematical Derivations

数学机理:<strong>chat template 的作用</strong>——把'角色 + 内容'的对话结构编码为<strong>模型可识别的 token 序列</strong>:通常用特殊 token 标记边界(如 <code><|system|></code>、<code><|user|></code>、<code><|assistant|></code>、<code><|end|></code>),并规定换行/空格等细节。<strong>为什么至关重要</strong>:(a) <strong>模型靠这些标记区分'谁在说话'</strong>——若无标记,模型无法知道该'回答'还是'续写';(b) <strong>训练-推理一致性</strong>——推理时必须用<strong>同一个</strong> template;若训练用 A、推理用 B(如少了一个换行、角色名不同),模型会'看不懂'、性能急剧下降(这是部署中最常见的坑之一);(c) <strong>system prompt 的支持</strong>——system 角色需在 template 中显式定义,否则模型无法学到'遵循系统指令'。<strong>多轮对话的构造</strong>:(1) <strong>历史拼接</strong>——把多轮历史按 template 拼接为一条长序列(<code>system + user1 + assistant1 + user2 + assistant2 + ...</code>);(2) <strong>损失位置</strong>——通常<strong>对所有 assistant 轮</strong>都算损失(教模型'每轮都回应');也有只算最后一轮的变体(更接近'基于历史回答当前问题');(3) <strong>上下文依赖</strong>——多轮样本能教模型'利用历史'(如指代消解、上下文延续),这是单轮数据学不到的;(4) <strong>长度控制</strong>——多轮历史可能很长,需截断策略(保留最近 N 轮 / 保留 system + 最近若干轮);(5) <strong>混入单轮数据</strong>——多轮数据(含大量历史)会让模型倾向于'长回答'或'依赖历史';故常混入单轮数据以平衡。<strong>其他要点</strong>——(a) <strong>特殊 token 的 embedding</strong>(新增 token 需训练);(b) <strong>EOS 与截断</strong>(正确放置结束标记,否则模型不会停);(c) <strong>训练数据的角色多样性</strong>(不同 system prompt、不同用户风格)。

🏭 Production Trade-offs

深度剖析与工程权衡:① <strong>'template 不一致'是部署第一大坑</strong>——训练与推理的 template 差一个空格/换行,就可能导致质量显著下降;故工程上应 (a) 把 template 存为<strong>配置文件</strong>、(b) 训练与推理<strong>共享同一份代码</strong>、(c) 加<strong>一致性测试</strong>(对比训练时与推理时的 token 序列)。② <strong>多轮 vs 单轮的配比</strong>——纯多轮数据会让模型过度依赖历史(在单轮场景表现差);纯单轮数据则学不到多轮能力。故需混合(如 50:50)。③ <strong>'损失算在哪些轮'的影响</strong>——算所有 assistant 轮会让模型'每轮都详细回应'(可能冗长);只算最后一轮更接近'问答'模式。选择取决于目标场景。④ <strong>与'长上下文'的关系</strong>——多轮对话天然需要长上下文能力;故多轮 SFT 常与长上下文训练配合(否则历史被截断)。⑤ <strong>system prompt 的训练</strong>——若训练数据中 system prompt 单一(或缺失),模型学不会'遵循多样化系统指令';故需<strong>多样的 system prompt</strong>(角色、风格、约束各不相同)。⑥ <strong>面试要点</strong>——被问'多轮对话怎么做 SFT',应给出'<strong>chat template 定义角色边界 + 历史拼接 + 损失落在 assistant 轮 + 截断策略 + 单轮/多轮混合</strong>',并强调'<strong>训练与推理必须同 template</strong>'这一部署要点;能指出'多轮与单轮需混合'是深度理解的标志。
⚠️ Common Interview Pitfalls
  • ✕
    训练与推理用不同的 chat template
  • ✕
    只用多轮数据(单轮场景表现差)
🎯 Interviewer Follow-ups
  • ?
    为什么 template 不一致会严重损害性能?
  • ?
    多轮训练与单轮训练如何混合?
📚

Associated Knowledge Base Guides & Mindmaps

Explore the comprehensive technical article, exam cards, and global architecture tree.

← PreviousM5-026: Instruction Tuning & Supervised Fine-Tuning: 解释 SFT 的学习率与训练轮数的选择。📋Back to BankNext →M5-028: Instruction Tuning & Supervised Fine-Tuning: 解释拒绝采样微调(RFT / STaR)的机制与作用。