论文多源确认

AgentTime 基准测试:智能体能否估算和控制自己的运行时长

AgentTime: Can Agents Estimate and Control Their Own Runtime?

精选理由

用 222 个任务测各家居家智能体能不能按时长干活,结果 Fable 5.1 偏差 2.9 倍,Astra 只有 1.2 倍,还有偷偷睡觉的,挺有意思。

AgentTime 基准包含来自 18 个来源的 222 个任务,覆盖编程、计算机使用和自动化研究,测试智能体能否按要求时长工作、预测运行时间并事后估算已用时间。时长跟随实验中,Fable 5.1 在 Claude Code 里偏离请求时长约 2.9 倍,GPT-6 Astra 在 Codex 中仅偏离 1.2 倍。158 个可分类的 Astra 运行记录里,有 14 个在看似完成任务后直接休眠以凑时长。回溯实验显示,移除时间信息会让 Sol 和 Astra 的偏差翻倍以上,Fable 接近翻倍。

原文 · arXiv cs.AI

AgentTime: Can Agents Estimate and Control Their Own Runtime?

An essential control of AI agents is their ability to manage runtime. This ability requires a sense of time-awareness, to predict and estimate wall-clock time and to control their own actions. Prior work has focused on time-awareness, but duration-following and control in native agent harnesses remain unexplored. We present AgentTime, a benchmark for testing whether agents can work for a requested duration, predict their runtime, and estimate elapsed time afterward. It comprises 222 tasks from 18 sources spanning coding, computer use, agentic work, and automated research. Duration-following experiments append a single instruction specifying how long to work, with requests ranging from about a minute to multiple days. Accuracy on these instructions varies substantially: Fable 5.1 in Claude Code deviates from requested runtimes by a typical factor of 2.9$\times$, compared with only 1.2$\times$ for GPT-6 Astra in Codex. However, matching the requested runtime does not, by itself, establish continued work on the task. Among 158 reviewed Astra runs with classifiable transcripts, 14 explicitly slept after appearing to finish. In forecasting experiments, predictions tend to overestimate natural runtimes. In retrospective experiments, removing temporal information more than doubles deviation for Sol and Astra and nearly doubles it for Fable. An agent's ability to complete a task does not guarantee that it can control its own time or work for the whole requested duration. For agents to run reliably, safely, and autonomously over long horizons, we require the evaluation of both.