用样条与项目反应理论重新估计 METR 的 AI 时间跨度
On the estimation and validity of AI time horizons---a statistical look at the METR plot
METR那个时间跨度指标被重新算了一遍:2到30分钟区间难度曲线几乎是平的,同样是10倍跨度,难度差别很大,附诊断图。
这篇论文在 228 个任务和 26 个 AI 上重新计算 METR 的 50% 时间跨度指标,用样条和项目反应理论替代了任务难度与人类用时对数呈线性关系的假设。拟合结果显示,人类用时从 2 到 30 分钟的区间内难度曲线接近平坦,其他区间接近线性,因此 3 分钟到 30 分钟的跨度提升比 30 分钟到 5 小时的提升容易得多,即使两者同样是 10 倍。作者给出的点估计在一组交叉验证的适当评分规则下表现更好,并提供了评估时间跨度构造效度的诊断图。论文建议在新增或扩长时间跨度基准时,应结合这些诊断图来解读指标。
On the estimation and validity of AI time horizons---a statistical look at the METR plot
METR's 50\% time horizon measures the human completion time of software tasks that an AI solves with 50\% probability, allowing AI capabilities to be expressed in interpretable units. On 228 tasks and 26 AIs, we recompute the time horizons using splines and item-response theory to relax the assumption that the AI difficulty of a task depends linearly on the log of human time. Our fitted spline can be interpreted as a function that \emph{converts} human time to AI difficulty; it is nearly flat in a region from 2--30 min but close to linear elsewhere. Hence, a time-horizon jump from 3 min to 30 min is much easier than one from 30 min to 5 hours despite the same multiplier of $10 \times$. Overall, we contribute time-horizon point estimates that perform better under a cross-validated suite of proper scoring rules, as well as diagnostic plots for assessing time horizons' construct validity. We suggest that time horizons be interpreted together with the diagnostic plots, especially as new time-horizon-based benchmarks are proposed or existing ones grow to include longer tasks.