Google发布Gemini 4 Argon模型
Google新模型Gemini 4 Argon在多项基准测试中击败对手,特别在编程和专业领域表现突出。
Google发布Gemini 4 Argon模型,在19项基准测试中13项排名第一。该模型在DeepSWE编程任务中得分为77.9%,高于Claude Opus 5.5的74.2%和GPT-6 Astra的74.1%。在Vals Index专业领域测试中得分为68.9%,输出限制提升至100万token。
A big day for Google! Gemini 4 Argon is here, taking first place on 13 of the 19 benchmarks Google published against GPT-6 Astra and Claude Opus 5.5. A few standout numbers: - DeepSWE (long real-world coding tasks) – 77.9%, vs. 74.2% for Opus 5.5 and 74.1% for Astra - Vals Index (real work in finance, law, tax, coding) – 68.9%, vs. 67.0% for Opus 5.5 and 63.1% for Astra - Output limit is 1M tokens, up from 64K Argon also leads on tests for vibe coding, long context, and reading charts and long videos. The release is limited to trusted cyber defenders in Google's Fairwind Program to start. 💬 3 🔄 0 ❤️ 8 👀 1749 📊 3 ⚡