论文

VISTA:多模态模型的视觉推理工具

VISTA: A Visual Harness for Reasoning in an Interactive World

精选理由

VISTA让Claude Opus在视觉游戏中表现超越人类,动作减少57.4%,多模态推理能力有突破。

研究人员推出VISTA,一种为多模态模型提供长程视觉能力的视觉工具。VISTA让模型能直接感知环境并保持无损视觉记忆,可主动检索过去观察结果。在ARC-AGI-3基准测试中,VISTA使Claude Opus 5.0的相对人类行动效率得分从40.68提升至满分100.00,模型完成全部25个公开游戏所用动作比首次参与的人类少57.4%。在另外三个视觉游戏和解谜基准测试中,VISTA也显著优于使用相同基础模型和简单工具的基线方法。

原文 · arXiv cs.AI

VISTA: A Visual Harness for Reasoning in an Interactive World

We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations in their original form. The model can actively retrieve these observations and reorganize its visual input as it reasons. On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00, with the model completing all 25 public games using 57.4% fewer actions than first-time human participants. VISTA's simple design also allows it to extend naturally to diverse visual environments with minimal adaptation. Across three additional benchmarks covering a diverse range of visual games and puzzles, it substantially outperforms baselines using the same underlying model with minimal harnesses. Our results highlight VISTA's potential as a general-purpose visual harness for advancing multimodal agents in complex visual environments.