论文

CVP 用目标亲和度 token 提升 3D 视觉语言模型表现

精选理由

UCSD 和 Lambda 的 CVP 方法让 3D 模型看清你问的是哪个物体,5 个基准全赢 Video-3D-LLM。

UC San Diego 与 Lambda 在 WACV 2026 提出 CVP 方法,解决 3D 视觉语言模型误判目标物体的问题。它引入目标亲和度 token 标记任务相关物体,并用 allocentric 网格提供全局上下文。对比 Video-3D-LLM,CVP 在 SQA3D EM 上从 58.6 提升到 62.3,在 Scan2Cap CIDEr 上从 83.8 提升到 90.5。该方法在 ScanQA、SQA3D、ScanRefer、Multi3DRefer、Scan2Cap 共 5 个基准上全部取得更好成绩。

图片来源 · Lambda
原文 · Lambda

Ask a 3D vision-language model what's near the table and in front of the curtain, and it might guess "sewing machine." The right answer is a tray rack.

CVP (UC San Diego + Lambda, WACV 2026) fixes this with a target-affinity token for task-relevant objects and an allocentric grid for global context.

Against Video-3D-LLM:

• SQA3D EM: 58.6 → 62.3 • Scan2Cap CIDEr: 83.8 → 90.5 • Better on all 5 benchmarks tested

Full results across ScanQA, SQA3D, ScanRefer, Multi3DRefer, and Scan2Cap, plus how the central/peripheral split works: https://t.co/O4YwNiES3r