模型多源确认精选

Shaped 审计 Zapier AutomationBench,修复 206 个验证器漏洞

精选理由

Shaped 给 Zapier 的 AutomationBench 做了一次验证器审计,用智能体找茬找出 206 个真实漏洞,近三成 Kimi K3 成绩因此改变,做评测的都该看看。

Shaped 对 Zapier 的 AutomationBench 中全部 600 个公开任务逐项检查验证器,用智能体提交貌似合理的错误答案来测试,人工复核确认 206 个真实漏洞并全部修复。团队重放 1235 次 Kimi K3 运行,其中 344 次(27.9%)成绩发生变化。验证器过严的任务上通过率从 18.8% 升至 43.8%,过宽的从 60.2% 降至 49.7%。修复后的版本以 AutomationBench Verified 发布。

原文 · elvis

We need more efforts like this. Every agent benchmark should audit its verifiers. Parsewave went through all 600 public tasks in Zapier's AutomationBench. Agents wrote realistic wrong answers to try to fool each verifier, and human review confirmed 206 real bugs. AutomationBench Verified fixed all 206. Regarding 1,235 Kimi K3 runs, the fixed verifiers changed 27.9% of the grades. I just started looking into this benchmark for some independent eval work I am doing, so this is good timing to see this audit. shaped @shaped AutomationBench Verified is out. We went through all 600 public tasks in @Zapier 's AutomationBench and checked every verifier. • Agents flagged 323 verifiers as suspicious, human review confirmed 206 real bugs, all 206 are fixed. • We replayed 1,235 Kimi K3 runs on the old and fixed verifiers and 344 of them (27.9%) got a different grade • Where verifiers were too strict, pass rate went from 18.8% to 43.8% • Where they were too lenient, it dropped from 60.2% to 49.7% Audit and dataset links below. 🔗 View Quoted Tweet 💬 9 🔄 2 ❤️ 14 👀 2286 📊 8 ⚡