测量优先的智能体排行榜审计框架:污染风险与评分器验证
Measurement-First Auditing of Agentic Leaderboards: Contamination Susceptibility, Matched-Control Re-evaluation, and Scorer Validation
一群研究者专门审计智能体排行榜的基准污染问题,用 GPT-4.1 和 DeepSeek-V4-Flash 在 SWE-bench 上做了对照实验,结论是很多污染指控证据不足,做评测的人值得看看这套框架。
论文提出一个测量优先的审计框架,把基准污染拆成训练期暴露、评估期检索和流水线/脚手架泄漏三条通道,并按 fail-closed 规则标记为 open、partial、closed 或 unknown。在 Holistic Agent Leaderboard(HAL)的 9 个配置、27 项通道评估中,没有任何一项被标记为 closed,但有 4 个配置确认了污染事件。团队随后在 SWE-bench Verified 上用同仓库匹配对照组设计检验 GPT-4.1 和 DeepSeek-V4-Flash,GPT-4.1 出现 +10.0 个百分点的 Top-3 基准相关差距,但配对自助法和仓库平衡分析给出的 95% 区间都包含零,差距结论为不确定。复现评分器未通过验证门槛:DeepSeek-V4-Flash 的两次触发在正确金标对比中均为假阳性,因此论文认为在缺乏来源证据和验证过的评分器时,不能支持更强的污染指控。
Measurement-First Auditing of Agentic Leaderboards: Contamination Susceptibility, Matched-Control Re-evaluation, and Scorer Validation
Agentic leaderboards increasingly evaluate systems on public benchmarks whose task statements and solution-bearing artifacts can remain accessible. We propose a measurement-first audit framework that distinguishes contamination claims according to the evidence required to support them. It separates three channels that require different evidence: training-time exposure, evaluation-time retrieval, and pipeline/scaffold leakage. Each channel is coded as open, partial, closed, or unknown under a fail-closed rule. Across nine Holistic Agent Leaderboard (HAL) configurations, none of the 27 channel assessments was coded closed, but incidents were confirmed in four configurations. We then apply the behavioral component of the framework to a reported file-localization gap on SWE-bench Verified, using an outcome-blind, same-repository matched-control design with symmetric prompt-leakage screening, paired and repository-aware uncertainty analyses, and scorer validation, evaluated on GPT-4.1 and DeepSeek-V4-Flash. Among the 100 pairs retained after symmetric screening and the pair-integrity exclusion, GPT-4.1 showed a $+10.0$-point pair-weighted Top-3 benchmark-associated gap, but the 95\% intervals from both the prespecified paired-bootstrap procedure and the post-hoc repository-balanced analysis included zero, leaving the benchmark-associated gap inconclusive. The reproduction scorer did not pass its validation gate: against consensus human labels, sufficient scorer sensitivity could not be established for either model, and both DeepSeek-V4-Flash firings on correct-gold comparisons were false positives. Without provenance evidence, appropriate controls, symmetric leakage screening, and validated scorers, stronger contamination claims are not warranted. The results do not establish training-data membership, contamination prevalence, or benchmark-induced score inflation.