Gemini 4 Argon 在文本和代码基准测试中领先
Google 新发布的 Gemini 4 Argon 在文本和代码基准测试中全面领先,性价比高达 8 美元/百万token。
Google DeepMind 发布的 Gemini 4 Argon (High) 在 Text Arena 基准测试中以 1525 分排名第一,在 Code Arena: WebDev 中以 1679 分排名第八。该模型在编程、复杂提示、指令遵循等五个子类别中排名第一,并在所有职业领域评估中领先。Gemini 4 Argon 比第二名 Claude Opus 4.6 高出 20 分,比 Google 上一版本 Gemini 3.8 Flash 提升了 96 分。
Big news: Gemini 4 Argon (High) by @GoogleDeepMind just landed #1 in Text Arena with 1525 pts, and #8 in Code Arena: WebDev with 1679 pts! This release has reshaped the Text Arena Pareto frontier with a blended $8/MToken! Gemini 4 Argon (High) is now the most cost efficient model, see its placement on Pareto frontier below. In the Text Arena, Gemini 4 Argon (High) ranks #1 in Coding, Hard Prompts, Instruction Following, Longer Query, and Creative Writing. It also leads every occupational domain evaluated, with additional #1 spots in English, Non-English, Chinese, and Russian. This model is +20 points above the #2 ranked Claude Opus 4.6 (High), and a huge leap from Google’s previous release, Gemini 3.8 Flash (High) at #11 ! In Code Arena: WebDev, Gemini 4 Argon (High) gained +96 points from Gemini 3.8 Flash (High), and went from #29 to #8 . Congrats to the @GoogleDeepMind team on this impressive frontier release! Google DeepMind @GoogleDeepMind Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program. 🔗 View Quoted Tweet 💬 34 🔄 55 ❤️ 762 👀 53994 📊 104 ⚡