M6-035M6: Multimodal & Generative ModelsVLM Training & EvaluationMedium
Mastery:

VLM Training & Evaluation: 解释 grounding 能力与它的训练数据。

📐 Mathematical Definition
grounding: text→region (x1,y1,x2,y2) or point;data: box-text pairs\text{grounding}:\ \text{text}\to\text{region }(x_1,y_1,x_2,y_2)\ \text{or point};\qquad \text{data}:\ \text{box-text pairs}
⚡ Executive Summary
Core Concept: grounding 指把语言指向图像的具体区域(输出框/点);训练需'区域-文本'标注数据(检测框、指代分割)。

📌 Key Takeaways

  • •
    能力:'图左上角的红色物体是什么'→ 定位并描述
  • •
    数据:检测框 + 文本描述、指代分割、视觉问答中的区域标注
  • •
    表示:坐标 token(量化到 bin)或特殊 token(如 <box>)

📐 Mathematical Derivations

数学机理:<strong>grounding 的定义</strong>——指模型能把<strong>语言指向图像的特定区域</strong>:(a) <strong>指代理解(referring expression)</strong>——'左边的那个红色杯子'→ 定位该物体;(b) <strong>区域描述</strong>——'这个框里是什么';(c) <strong>输出定位</strong>——模型生成坐标(框/点)。<strong>与'整体理解'的区别</strong>——整体理解只需'描述大致内容';grounding 要求<strong>精确的空间对应</strong>(模型必须知道'哪个 patch 对应哪个物体')。<strong>训练数据</strong>——(a) <strong>检测数据</strong>(目标检测的框 + 类别,如 COCO Detection);(b) <strong>指代分割(referring segmentation)</strong>(文本 + 像素级掩码,如 RefCOCO/RefCOCOg);(c) <strong>区域描述(region caption)</strong>(框 + 描述,如 Visual Genome);(d) <strong>图文对的隐式 grounding</strong>(弱监督:整句描述 → 隐式学到区域对应);(e) <strong>合成数据</strong>(程序生成'框 + 文本'对,规模大、质量可控)。<strong>坐标的 token 化</strong>——模型需'输出坐标',但坐标是连续值;常用方法:(a) <strong>量化到 bin</strong>——把坐标归一化到 [0,1] 后离散化为 N 个 bin(如 1000 个),作为<strong>特殊 token</strong>(如 <code><loc_512></code>)输出;(b) <strong>文本化坐标</strong>——直接输出数字字符串(如 <code>[0.23, 0.45, 0.67, 0.89]</code>);(c) <strong>专门的位置 token</strong>(如 <code><box></code> 包裹)。<strong>Qwen2-VL 的做法</strong>——用 <code><|box_start|>(x1,y1),(x2,y2)<|box_end|></code> 的格式,坐标量化到 0~1000 的整数。<strong>为什么 grounding 难</strong>——(a) <strong>空间精确性</strong>——需要模型保留 patch 的空间位置信息(故需 2D RoPE 或位置嵌入);(b) <strong>数据稀缺</strong>——区域级标注比'图文对'贵得多;(c) <strong>与语言的对齐</strong>——需把'左侧/上方'等空间词与坐标对应;(d) <strong>高分辨率</strong>——小物体需高分辨率才能定位(见动态分辨率)。<strong>应用</strong>——(a) <strong>视觉问答中的精确定位</strong>('图中的第二个人在做什么');(b) <strong>UI 操作</strong>(点击某按钮 → 输出坐标,见 computer use);(c) <strong>图像编辑</strong>('把左边的树去掉');(d) <strong>机器人</strong>('拿起桌上的杯子')。<strong>评估</strong>——(a) <strong>RefCOCO/RefCOCO+/RefCOCOg</strong>(指代分割/定位的准确率);(b) <strong>Visual Genome 区域描述</strong>;(c) <strong>定位精度</strong>(IoU);(d) <strong>UI 定位基准</strong>(如 ScreenSpot)。<strong>提升手段</strong>——(a) <strong>专门的 grounding 数据微调</strong>;(b) <strong>高分辨率 + 2D 位置编码</strong>;(c) <strong>显式的坐标监督</strong>(训练时要求输出坐标);(d) <strong>与检测模型结合</strong>(用检测器提供候选区域)。

🏭 Production Trade-offs

深度剖析与工程权衡:① <strong>'grounding 需要空间信息保留'</strong>——它是'2D RoPE / 位置嵌入'的直接动机之一(若位置信息丢失,无法定位);故 grounding 与位置编码设计强相关。② <strong>'坐标 token 化'是工程细节但关键</strong>——量化到 bin 是最常用方案(兼顾精度与 token 效率);bin 数需权衡(太少则不精确、太多则 token 变长)。③ <strong>'区域标注贵'</strong>——这是 grounding 能力的瓶颈;故常用 (a) 弱监督(图文对的隐式 grounding)、(b) 合成数据、(c) 用检测器自动标注。④ <strong>'UI 定位是 grounding 的重要应用'</strong>——computer use Agent 需要'点击某按钮'→ 输出坐标;故有专门的 UI 定位基准(ScreenSpot)与训练数据。⑤ <strong>'与检测模型的分工'</strong>——专用检测器在'标准类别'上更准;VLM 的 grounding 优势在'开放词汇 + 语言理解'(如'左边第三个红色物体')。⑥ <strong>面试要点</strong>——被问'grounding 是什么',应给出'<strong>语言→区域的精确定位 + 训练数据(检测框/指代分割/区域描述)+ 坐标 token 化(量化到 bin)</strong>'与'<strong>空间信息保留(2D RoPE)与数据稀缺是难点</strong>';能指出'UI 定位是重要应用'是深度理解的标志。
⚠️ Common Interview Pitfalls
  • ✕
    以为整体描述能力等于 grounding 能力
  • ✕
    坐标不做量化直接输出浮点数(token 效率低)
🎯 Interviewer Follow-ups
  • ?
    为什么 grounding 需要专门数据?
  • ?
    坐标如何 token 化?
📚

Associated Knowledge Base Guides & Mindmaps

Explore the comprehensive technical article, exam cards, and global architecture tree.

← PreviousM6-034: VLM Training & Evaluation: 解释 VLM 幻觉的类型与缓解。📋Back to BankNext →M6-036: VLM Training & Evaluation: 解释 VLM 在多图与视频上的评估难点。