VeriFine:随策略共同进化的具身推理验证框架
VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning
论文讲了个有意思的思路:策略会不断冒出新失败模式,验证器也得跟着一起进化。适合做具身智能体和自我改进研究的同学看。
arXiv 论文提出 VeriFine,一个让验证器与策略、训练课程共同进化的 agent harness 框架。Policy Improvement Loop 用 rubric judge 诊断反复出现的失败并构建自适应课程,Judge Improvement Loop 在策略进步停滞时挑选关键失败案例向人类查询,通过 coactive calibration 校准验证器。实验覆盖驾驶与机器人导航两类任务,在强化学习和监督微调下实现策略与验证器能力的持续提升。
VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning
Self-improving policies continually expose new failure patterns, changing what their judges must be able to verify. However, current fixed judges constrain both optimization feedback and the discovery of useful training examples, limiting further self-improvement. This challenge is even more acute in embodied reasoning, where reliable evaluation must account for spatial grounding, causal reasoning, and safety-aware decision-making. We introduce VeriFine, an agent harness framework that scales verification through the co-evolution of the policy, training curriculum, and judge. The Policy Improvement Loop uses a rubric judge to diagnose recurring failures, construct an adaptive curriculum, and optimize the policy. When progress plateaus and verification becomes a bottleneck, the Judge Improvement Loop selectively queries human guidance on informative failure cases and refines the judge through coactive calibration, in which humans and agents resolve disagreements and converge toward the objective rubric of physical reasoning. The revised judge then guides the next stage of data selection and policy optimization. Experiments on driving and robot navigation tasks demonstrate continuous self-improvement in both policy and judge capability across reinforcement and supervised fine-tuning. These results show how scaling verification supports continuous self-improvement as policy failure patterns evolve.