Gemini 4 在AAI指数表现优异
Gemini 4在多个基准测试中表现优异,特别是低幻觉率值得关注。
Gemini 4 Argon在Artificial Analysis的Intelligence Index上获得53分,与GPT-6 Astra持平,高于GPT-6.1 Sol的52分。该模型在AutomationBench-AA基准测试中排名第一。在AA-Omniscience测试中,Argon的幻觉率仅为15%,远低于Astra的51%,但答案准确率为50%,低于Astra的63%。
Even on AAI index Gemini 4 looks really really good! Fable-Class Model. Gemini 4 Argon matches GPT-6 Astra on Artificial Analysis’s Intelligence Index.
53 points, versus Astra’s 53 and GPT-6.1 Sol’s 52. It also takes #1 on AutomationBench-AA.
On AA-Omniscience, Argon hallucinates far less: 15% versus Astra’s 51%. It is more willing to admit uncertainty, though its answer accuracy is lower: 50% versus 63%.
Yep, looks like a very good release. Kudos @GeminiApp