模型多源确认83°

Claude Sonnet 5.5 在 Code Arena 排名第四

Real-world results are in for Claude Sonnet 5.5 (High) by @AnthropicAI. It just landed #4 in Code Ar...

精选理由

Anthropic 发布的 Claude Sonnet 5.5 性价比极高,在多个关键领域排名从第 30 多位跃升至第 4 位。

Claude Sonnet 5.5 (High) 在 Code Arena: WebDev 基准测试中获得 1699 分,排名第 4 位。该模型比排名第 3 的 Claude Fable 5.1 (Max) 和排名第 2 的 GPT-6 Astra (Max) 便宜 80%,每百万代币成本为 8 美元。相比 Sonnet 5 (High),该模型性能提升了 159 分,排名从第 37 位跃升至第 4 位。

原文 · lmarena.ai

Real-world results are in for Claude Sonnet 5.5 (High) by @AnthropicAI. It just landed #4 in Code Ar...

Real-world results are in for Claude Sonnet 5.5 (High) by @AnthropicAI . It just landed #4 in Code Arena: WebDev with 1699 pts, and has reshaped the Pareto frontier with its cost efficiency! Claude Sonnet 5.5 (High) delivers nearly top performance at a blended $8 per Mtoken, reshaping the Pareto frontier! This model is 80% cheaper than both Claude Fable 5.1 (Max) in the #3 spot overall, and GPT-6 Astra (Max) at #2 . See Pareto placement below. Overall, Claude Sonnet 5.5 (High) is a +159 pt improvement from Sonnet 5 (High) at #37 with 1540 pts. This gain compared to its previous variant also shows up across these key domains so far: - Reference-Based Design: #38 → #4 - Simulations: #37 → #4 - Gaming: #36 → #4 Congrats to @AnthropicAI on this release! Claude @claudeai Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work. Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 2 🔄 1 ❤️ 4 👀 1315 📊 2 ⚡