视觉语言模型空间推理新方法GPD
Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models
清华团队提出GPD方法,让视觉语言模型通过3D几何特权提升空间推理能力,无需改变模型架构。
研究人员提出GPD方法,通过几何特权的蒸馏技术提升视觉语言模型的空间推理能力。该方法在4B参数模型上达到57.1的VSI-Bench分数,在多个空间推理基准测试中平均得分为37.6。GPD将深度、语义和鸟瞰图线索作为特权信息,仅对错误轨迹应用KL散度,保持模型仅使用RGB输入。
Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models
Spatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence. Existing remedies either inject 3D into the model at inference, paying architecture and latency costs, or train with outcome rewards that supervise only the final answer. Spatial errors originate in perception: a misjudged depth or direction can be corrected only by the scene's true geometry, which the 3D-scanned sources of spatial training corpora already provide. We propose GPD (Geometry-Privileged Distillation), which makes geometric evidence the privilege in on-policy self-distillation (OPSD). For each question, depth, semantic, and bird's-eye-view (BEV) cues are rendered as compact text and routed to the teacher alongside the reference answer; a privileged KL, applied only to incorrect trajectories, augments GRPO, and the deployed model remains RGB-only. On the 4B backbone, GPD achieves 57.1 on VSI-Bench and 37.6 average across MindCube, SPARBench, MMSI-Bench, and ViewSpatial, outperforming both GRPO and answer-privileged OPSD across spatial reasoning benchmarks. Ablations confirm the complementarity of 3D and answer privilege, the advantage of question-conditioned routing over full-context injection, and the benefit of restricting distillation to incorrect trajectories.