Shaped 审计 Zapier AutomationBench:600 个任务中发现 206 个验证器 bug
Shaped 把 Zapier 的 AutomationBench 全查了一遍,206 个验证器有 bug,修完后 Kimi K3 近三成评分都变了,做评测的都该看看。
Shaped 团队对 Zapier AutomationBench 的全部 600 个公开任务逐一检查验证器,让智能体生成看似合理的错误答案来攻击每个验证器,经人工复核确认 206 个真实 bug 并全部修复。团队用修复后的验证器重跑 1,235 次 Kimi K3 运行,27.9%(344 次)的评分发生改变。验证器过严的任务通过率从 18.8% 升至 43.8%,过松的任务从 60.2% 降至 49.7%。修复后的版本以 AutomationBench Verified 发布,附带审计方法和数据集。
We need more efforts like this. Every agent benchmark should audit its verifiers. Parsave went through all 600 public tasks in Zapier's AutomationBench. Agents wrote realistic wrong answers to try to fool each verifier, and human review confirmed 206 real bugs. AutomationBench Verified fixed all 206. Regrading 1,235 Kimi K3 runs with the fixed verifiers changed 27.9% of the grades. I just started looking into this benchmark for some independent eval work I am doing, so this is good timing to see this audit. shaped @shaped AutomationBench Verified is out. We went through all 600 public tasks in @Zapier 's AutomationBench and checked every verifier. • Agents flagged 323 verifiers as suspicious, human review confirmed 206 real bugs, all 206 are fixed. • We replayed 1,235 Kimi K3 runs on the old and fixed verifiers and 344 of them (27.9%) got a different grade • Where verifiers were too strict, pass rate went from 18.8% to 43.8% • Where they were too lenient, it dropped from 60.2% to 49.7% Audit and dataset links below. 🔗 View Quoted Tweet 💬 2 🔄 0 ❤️ 6 👀 1129 📊 2 ⚡