GPT-6 Astra在ARC-AGI-3测试中首次超越人类效率,让AI专家Chollet提前了AGI预测时间。
OpenAI的GPT-6 Astra在基准测试中表现不一。Epoch AI将其评为169分,排名第一;而Artificial Analysis则认为其表现与前任相当,落后于Claude Fable 5.1。最引人注目的是在ARC-AGI-3测试中,Astra首次以超越普通人类的效率完成测试。ARC Prize负责人François Chollet虽不认为这已证明AGI实现,但他认为进步速度比预期快一倍,并提前了AGI预测时间。
Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward
OpenAI's GPT-6 Astra is drawing contradictory benchmark verdicts. Epoch AI puts it out in front with 169 points, while Artificial Analysis rates it no better than its predecessor and behind Claude Fable 5.1. The biggest surprise comes from ARC-AGI-3, where Astra works more efficiently than the average human for the first time. ARC Prize chief François Chollet doesn't call this proof of AGI, but he does see the progress running "twice as fast" as he expected, and he's moving up his AGI forecast. The article Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward appeared first on The Decoder .