论文多源确认精选

Salesforce CLIFT 方法让 31B Gemma-4 网页智能体在 WebArena 超过 Gemini 3 Flash

精选理由

Salesforce 的新论文 CLIFT,31B 开源模型打网页任务居然赢了 Gemini 3 Flash,关键是部署时不用裁判模型,省成本。

Salesforce AI Research 提出 CLIFT 方法,让 31B 开源 Gemma-4 网页智能体在 9 个应用的 WebArena Infinity 上拿到 74.6%,高于使用 browser use 的 Gemini 3 Flash 的 70.1%。方法核心是让智能体回答关于自身 rollout 的验证问题,再用保形认证器筛选与训练时裁判一致的问题并加权加入每步奖励。测试时只用冻结题库在贪婪 rollout 和少量重试之间选择,无需外部裁判。该智能体比基础模型提升 12.8 分,在 9 个应用中赢下 7 个,题库还能迁移到 GPT-5.5(VisualWebArena)和 Online Mind2Web 的实时网页智能体上。

原文 · DAIR.AI

Another great paper from Salesforce AI Research.

The finding is that a 31B open Gemma-4 web agent scores 74.6% on the 9-app WebArena Infinity set, above Gemini 3 Flash with browser use at 70.1%.

They got there without calling a frontier judge at every step or at deployment.

CLIFT has the agent answer verification questions about its own rollouts.

A conformal certifier keeps only the questions whose answers agree with a training-time judge, weights them by how much they can be trusted, and adds the result to per-step rewards.

At test time, the same frozen question bank picks between a greedy rollout and a few retries, with no external judge.

The trained agent improves 12.8 points over its base model and wins 7 of 9 apps. The question bank also transfers to GPT-5.5 at test time on VisualWebArena, and a translated bank improves a live-web agent on Online Mind2Web without any training on that benchmark.

Paper: https://t.co/kzHeT4CYhW

Chat with Paper: https://t.co/SoMbjIxAuK

  • Sundar Pichai10-06 16:03原文
  • Google DeepMind10-06 16:04原文