模型83°

Gemini 4 Argon (High)在Agent Arena排名第8

More big news from @GoogleDeepMind: Gemini 4 Argon (High) is #8 in Agent Arena with a +7.92% net imp...

精选理由

GoogleDeepMind的Gemini 4 Argon (High)在多个基准测试中表现优异,成本效益高,特别是在文本生成和网页开发领域。

GoogleDeepMind的Gemini 4 Argon (High)在Agent Arena中排名第8,净改进分数为+7.92%。相比Gemini 3.8 Flash (High),提升了4.96个百分点。该模型在Steerability指标上排名第一,达到+15.88%。在Text Arena中,Gemini 4 Argon (High)以1525分排名第一,在Code Arena: WebDev中以1679分排名第8。

原文 · lmarena.ai

More big news from @GoogleDeepMind: Gemini 4 Argon (High) is #8 in Agent Arena with a +7.92% net imp...

More big news from @GoogleDeepMind : Gemini 4 Argon (High) is #8 in Agent Arena with a +7.92% net improvement score, and has reshaped the Pareto frontier with a $0.62 cost per task! See its placement below. Gemini 4 Argon (High) is a 4.96 percentage point improvement over Gemini 3.8 Flash (High), at #19 with +2.96% net improvement. By key signals, Gemini 4 Argon (High) stands out in: - #1 in Steerability with +15.88% (the model’s ability to course-correct when you push back) - #2 in Confirmed Success with +14.15% (explicit user feedback that the task worked) - #4 in Praise vs Complaint with +27.74% (implicit sentiment in user reactions) By category Gemini 4 Argon (High) is especially strong in Chat, landing at #3 with +11.58% net improvement. With 3k real-world agentic sessions so far, this score is preliminary. Stay tuned as more traces come in from our global community of users. Congrats to the @GoogleDeepMind team on this release! Arena.ai @arena Big news: Gemini 4 Argon (High) by @GoogleDeepMind just landed #1 in Text Arena with 1525 pts, and #8 in Code Arena: WebDev with 1679 pts! This release has reshaped the Text Arena Pareto frontier with a blended $8/MToken! Gemini 4 Argon (High) is now the most cost efficient model, see its placement on Pareto frontier below. In the Text Arena, Gemini 4 Argon (High) ranks #1 in Coding, Hard Prompts, Instruction Following, Longer Query, and Creative Writing. It also leads every occupational domain evaluated, with additional #1 spots in English, Non-English, Chinese, and Russian. This model is +20 points above the #2 ranked Claude Opus 4.6 (High), and a huge leap from Google’s previous release, Gemini 3.8 Flash (High) at #11 ! In Code Arena: WebDev, Gemini 4 Argon (High) gained +96 points from Gemini 3.8 Flash (High), and went from #29 to #8 . Congrats to the @GoogleDeepMind team on this impressive frontier release! 🔗 View Quoted Tweet 💬 4 🔄 8 ❤️ 98 👀 8713 📊 13 ⚡