Timeline-Bench:56 个真实视频剪辑任务,最强智能体只完成 26.8%
Timeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut
让 AI 智能体真正剪片子,结果最强的 GPT-6 Astra 也只完成了不到三成,论文还开源了全部任务,可以自己上手测。
arXiv 论文 Timeline-Bench 收录 56 个真实视频剪辑任务,要求智能体从原始素材产出成片,覆盖对白选段、访谈叙事和广告剪辑。质量测试基于 43 位剪辑师的 2,582 次盲评校准。论文评测了 16 个由 GPT-6 Astra、Claude Code、Codex、OpenCode 等驱动的智能体,最优成绩是 GPT-6 Astra 配合 Codex 完成 56 个任务中的 15 个(26.8%),平均完成率 14.0%。771 次失败运行中有 562 次只败在质量测试:智能体通过静帧和转写稿理解素材,无法把控成片工艺。任务、验证器和逐次运行结果已在官网开源。
Timeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut
AI agents increasingly carry out long-horizon professional work, but their evaluations rarely require a finished creative deliverable. To this end, we introduce Timeline-Bench, a benchmark of 56 real video-editing tasks, each asking an agent to turn raw production material into a finished video. Tasks range from selecting dialog takes and shaping interview footage into a story to cutting commercials from product shots, voiceovers and graphics. Every task provides a brief, source assets, a container and a set of tests. A task is resolved when the output passes every test. The tests check the delivery format, the content and the brief's explicit requirements, and include a quality test calibrated on 2,582 blind judgments by 43 video editors. We evaluate 16 agents that pair frontier models with coding-agent harnesses such as Codex, Claude Code and OpenCode. The best, GPT-6 Astra in Codex with curated editorial guidance, resolves only 15 of the 56 tasks (26.8%), and the average agent resolves 14.0%. Human editors prefer the reference edit in 83.5% of judgments. Most unresolved runs (562 of 771) fail only the quality test: agents perceive footage through stills and transcripts and check their renders for defects, not craft. We release the tasks, verifier and per-run results at https://timelinebench.tensortest.com.