论文

回顾性世界建模提升VLM智能体推理能力

Beyond Prediction: Steering VLM Agents with Retrospective World Modeling

精选理由

这篇论文提出让VLM智能体不仅能预测未来,还能回溯过去判断因果一致性,解决了物理行为不连贯的问题。

研究人员提出回顾性世界建模(Retrospective World Modeling)新范式,使智能体能反向推理状态转换的因果一致性。该方法通过估计动作归因分布P(â_t|s_t,s_{t+1}),构建自一致性奖励(SCR)信号。实验表明,该方法在多样化智能体任务中显著提升了策略的鲁棒性和泛化能力,优于仅依赖前瞻性预测的基线模型。

原文 · arXiv cs.AI

Beyond Prediction: Steering VLM Agents with Retrospective World Modeling

Equipping VLM agents with world modeling capabilities has shown strong potential for complex reasoning and long-horizon planning, while reducing the dependence of policy learning on costly real-world interactions. Existing methods mainly rely on prospective simulation to predict the consequences of candidate actions. However, this forward-only paradigm focuses on what will happen next and provides limited constraints for verifying whether an action is causally consistent with the observed state transition, which can lead to plausible-looking but physically incoherent behaviors. In this paper, we challenge the view of world modeling as only prospective prediction and introduce Retrospective World Modeling, a new agent learning paradigm that enables agents to reason backward by estimating the retrospective attribution distribution $P(\hat{a}{t}|s_t, s{t+1})$ for the action that most likely caused a given transition. Based on this capability, we formulate the Self-Consistency Reward (SCR), an intrinsic signal that measures the probabilistic consistency between the policy action and the retrospective explanation. Integrating SCR into reinforcement learning provides dense transition-level feedback and steers agents toward behaviors that are both task-effective and physically grounded. Extensive experiments across diverse agentic tasks show that our method substantially improves policy robustness and generalization over prospective-only world modeling baselines.