EviRover:强化智能体感知能力
EviRover: Reinforcing Agentic Perception Beyond a Glance
EviRover通过交互解决单次观察不足的感知问题,在多个基准测试上大幅提升性能,已开源所有资源。
研究人员提出EviRover,首个通过交互解决感知查询的智能体模型。该模型使用监督微调和强化学习训练,在EviLens基准测试上比骨干模型平均提高30分。4B参数的EviRover性能接近先进专有模型,在BrowseComp-VL等多模态基准上也有15点提升。团队开源了所有代码、模型和数据集。
EviRover: Reinforcing Agentic Perception Beyond a Glance
Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumption often fails in real-world scenarios that hinge on fine-grained visual details or require knowledge-intensive and up-to-date information. We term such cases \textit{perception under insufficient evidence} and formulate perception as an agentic process that can obtain information beyond a single glance. To address the absence of data for this setting, we design two dedicated data generation pipelines, yielding EviRover-SFT-5K and EviRover-RL-12K for training. We further construct EviLens, a human-verified benchmark comprising 688 instances across five perception categories. Building on these data, we present EviRover, to our knowledge the first perception agent explicitly trained to resolve perceptual queries through interaction, using supervised fine-tuning followed by agentic reinforcement learning. Experiments show that the 4B EviRover outperforms its backbone by 30 points on average on EviLens, reaching performance comparable to advanced proprietary models. The gains transfer beyond EviLens to WebEyes, conventional perception benchmarks, and general multimodal benchmarks, including a 15-point improvement on BrowseComp-VL. All code, models, and data are released.