OpenAI的Astra在多个基准测试中创纪录,但ECI分数仍在趋势线内,GPT-6距离AGI仍有差距。
Astra模型在ECI基准测试中获得169分,创下新纪录,较之前最佳成绩163分有显著提升。该模型还在数学、持续学习和游戏谜题基准测试中创造了新纪录。在长程编程基准MirrorCode测试中,Astra表现介于Opus 4.7和Fable 5之间。OpenAI提供了Astra的预发布访问权限供测试。
Astra is a genuine improvement but not even off trend. So much for GPT-6 being AGI, folks. (i think...
Astra is a genuine improvement but not even off trend. So much for GPT-6 being AGI, folks. (i think we can safely assume that real AGI would be significantly above the trend line,) Epoch AI @EpochAIResearch GPT-6 Astra has set a new ECI record, with a score of 169. This is a substantial jump from the prior best (163), but is within our uncertainty range for the reasoning-era ECI trend. Astra also set new records on our math, continual learning, and game-puzzles benchmarks. On our long-horizon coding benchmark, MirrorCode, Astra ranks between Opus 4.7 and Fable 5. OpenAI gave us pre-release access to test Astra. Charts and more details for Astra’s individual benchmark results in the thread. 🔗 View Quoted Tweet 💬 1 🔄 2 ❤️ 6 👀 1107 📊 2 ⚡