将参考图像作为"原生词汇"直接嵌入文本指令的对应语义位置,解决多图像复杂交织指令下的对象绑定失效问题。
[Image1][Image2][Image3] + "A robot in image 1 holds a vase from image 2..."| 贡献 | 描述 |
|---|---|
| INSET 模型 | 将图像嵌入为句子中的原生词汇,精确对象绑定 |
| 可扩展数据引擎 | 从图像+视频数据自动构建 1500 万高质量交织样本 |
| InterleaveBench | 专为复杂多图任务设计的评测基准 |
核心思路:Native Interleaved Formulation
旧范式(间接引用):
[Image1][Image2][Image3] + "A robot in image 1 holds a vase from image 2..."
INSET 新范式(原生嵌入):
"A [Image1_tokens] robot holds a [Image2_tokens] flower vase for a [Image3_tokens] dog."
模型架构