M6-070M6: Multimodal & Generative ModelsFlow Matching & Rectified FlowMedium
Mastery:

Flow Matching & Rectified Flow: 解释 Flow Matching 与 score matching 的关系。

📐 Mathematical Definition
vmarginal=E[vcond∣xt];∇log⁡pt=E[∇log⁡p(xt∣x0)∣xt]v_{\text{marginal}}=\mathbb{E}[v_{\text{cond}}|x_t];\qquad \nabla\log p_t=\mathbb{E}[\nabla\log p(x_t|x_0)|x_t]
⚡ Executive Summary
Core Concept: 两者都是'用可计算的条件量回归不可计算的边缘量':FM 回归条件速度、score matching 回归条件 score;扩散是两者的桥梁。

📌 Key Takeaways

  • •
    结构相同:条件量可采样、边缘量不可算、条件期望 = 边缘量
  • •
    FM 回归'速度',score matching 回归'score'
  • •
    两者可互相转换(速度与 score 有线性关系)

📐 Mathematical Derivations

数学机理:<strong>相同的数学结构</strong>——(1) <strong>score matching</strong>——目标 ∇_x log p_t(x)(边缘 score,不可计算);用<strong>条件 score</strong> ∇_x log p(x_t|x_0)(可计算)回归;关键恒等式 <strong>E[∇log p(x_t|x_0)|x_t]=∇log p_t(x_t)</strong>(条件期望 = 边缘量)。(2) <strong>Flow Matching</strong>——目标 v_marginal(边缘速度,不可计算);用<strong>条件速度</strong> v_cond(可计算)回归;关键恒等式 <strong>E[v_cond|x_t]=v_marginal(x_t)</strong>。<strong>两者都是'条件期望 = 边缘量'的应用</strong>;故训练都是'回归可计算的条件量'。<strong>互相转换</strong>——对<strong>扩散路径</strong>(x_t=√ᾱ_t x_0+√(1−ᾱ_t)ε),可证明:<strong>v = f(x,t) − ½g²(t)·∇log p_t</strong>(来自 SDE/ODE 的关系),即'速度场 = 漂移项 − ½扩散系数² × score';故 <strong>(a) 若知道 score,可算速度</strong>;(b) 若知道速度,可算 score。<strong>这解释了'为什么扩散与 FM 是同一框架'</strong>——它们是同一个 SDE/ODE 的两种'参数化'。<strong>训练难度的差异</strong>——(a) <strong>score matching</strong> 需'回归 score'(其尺度随 t 变化剧烈,需参数化技巧如 ε/v-prediction);(b) <strong>FM</strong> 直接回归速度(对直线路径是常数 x_1−x_0,<strong>尺度均匀</strong>)——故 FM 的训练目标<strong>更简单、方差更小</strong>(这是它被认为'更易训'的原因之一)。(c) 但对'扩散路径',FM 的速度与 score 有线性关系,故<strong>等价</strong>(训练难度相当)。<strong>实践意义</strong>——(a) <strong>统一视角</strong>——可用同一套代码/理论处理扩散与 FM(只是'路径'不同);(b) <strong>技术迁移</strong>——score matching 的采样器(SDE/ODE 求解器)、参数化技巧可直接用于 FM;(c) <strong>引导</strong>——CFG 在 FM 中也可表达(用条件与非条件速度的组合,形式与扩散相同)。<strong>其他相关</strong>——(a) <strong>denoising score matching(DSM)</strong>——扩散损失的另一种推导;(b) <strong>SDE/ODE 的统一框架</strong>(Song 等);(c) <strong>随机插值(stochastic interpolants)</strong>——更一般的框架,包含扩散与 FM 为特例。<strong>实践</strong>——(a) <strong>扩散</strong>——用 ε/v-prediction + score matching 的等价损失;(b) <strong>FM</strong>——直接回归速度;(c) 两者可互换(对同一路径);(d) 现代模型(SD3、Flux)用 FM(因为训练目标更简单)。

🏭 Production Trade-offs

深度剖析与工程权衡:① <strong>'条件期望 = 边缘量'是统一的核心</strong>——面试中能指出'FM 与 score matching 结构相同'是深度理解的标志。② <strong>'速度与 score 的线性关系'是桥梁</strong>——v=f−½g²∇log p;它解释了'为什么两种参数化等价'。③ <strong>'FM 训练目标更简单'</strong>——对直线路径,v_cond=x_1−x_0 是常数(尺度均匀),而 score 的尺度随 t 剧变;故 FM 更易训(无需复杂的参数化技巧)。④ <strong>'技术可迁移'</strong>——采样器、引导、量化等都可在两个框架间迁移;这是'统一视角'的实用价值。⑤ <strong>'随机插值'是更一般的框架</strong>——它包含扩散与 FM 为特例;理解它能把握整个'生成模型谱系'。⑥ <strong>面试要点</strong>——被问'FM 与 score matching 的关系',应给出'<strong>结构相同(条件期望 = 边缘量)+ 速度与 score 线性相关(v=f−½g²∇log p)+ 扩散是桥梁</strong>'与'<strong>FM 训练目标更简单</strong>';能写出 v 与 score 的关系式是深度理解的标志。
⚠️ Common Interview Pitfalls
  • ✕
    以为 FM 与扩散是两套无关的框架
  • ✕
    忽略'条件期望 = 边缘量'这一共同结构
🎯 Interviewer Follow-ups
  • ?
    为什么两者可以互相转换?
  • ?
    哪个训练更简单?
📚

Associated Knowledge Base Guides & Mindmaps

Explore the comprehensive technical article, exam cards, and global architecture tree.

← PreviousM6-069: Flow Matching & Rectified Flow: 解释条件流匹配与边缘流匹配的差异。📋Back to BankNext →M6-071: Flow Matching & Rectified Flow: 解释 Flow Matching 中的路径选择与最优传输。