论文

DAYJOB专业工作基准发布

DAYJOB: A Benchmark for Long-Horizon Professional Work

精选理由

研究人员发布了DAYJOB基准,测试了13家开发者的30种模型配置,发现专业领域AI表现仍有巨大提升空间。

DAYJOB基准包含130个专业任务,其中医疗领域50个,金融领域80个。每个任务平均需要专业人员13.6小时(医疗)和16.6小时(金融)完成。测试显示,Claude Opus 5.5表现最佳,通过率24.7%(医疗)和23.9%(金融),而中位数模型仅通过0.6%和2.5%。研究还发现AI代理会接受与记录相矛盾的前提并错误传递输入数据。

原文 · arXiv cs.AI

DAYJOB: A Benchmark for Long-Horizon Professional Work

Professional work often starts with a brief request that leaves the professional to work out what is needed, which documents matter, and whether the request's premise holds. We introduce DAYJOB, a benchmark of 130 tasks built by professionals in healthcare (50) and finance (80). The tasks are estimated to take a professional 13.6 hours on average in healthcare and 16.6 in finance. Each task is a containerized Harbor environment with an expert rubric of binary criteria (median 47.5 and 57.5 per task) that an agentic judge applies to the delivered files, and an attempt passes only if it meets every criterion. Across 30 model configurations from 13 developers, the strongest, Claude Opus 5.5, passes 24.7% of healthcare and 23.9% of finance attempts, and the median configuration passes 0.6% and 2.5%. In case studies, agents accept premises that the record contradicts and carry wrong inputs through otherwise consistent analyses. We release all healthcare tasks, 50 of the 80 finance tasks, the evaluation harness, and the leaderboard.