Epoch AI 推出 InnovationEval 基准,测试 AI 能否做出后训练创新
Epoch AI 做了个叫 InnovationEval 的基准,专门测 AI 智能体能不能自己搞出像样的后训练创新,目前结果挺拉胯的,想追自动化研究进展的可以看看。
Epoch AI 构建了 InnovationEval 基准,用于检验 AI 智能体能否做出与近期已发表进展同等量级的 post-training 创新。该基准衡量的是 AI 产出创新结果的幅度,而非仅仅复现论文。目前测试结果显示,各智能体的表现仍不理想,与人类研究者已发表的成果存在明显差距。这为评估自动化 AI 研究的实际进展提供了一个可量化的标尺。
AI developers aim to create an automated AI researcher. How close are they?
To find out, we built InnovationEval, which tests whether AI can produce post-training innovations comparable in magnitude to a recently published advance. So far, agents’ results are underwhelming. https://t.co/EGnTOL0klR
- rohanpaul_ai10-05 19:04原文