Claude Opus 5.5 在 Agent Arena 排名第二
Claude Opus 5.5 (High) from @AnthropicAI just entered Agent Arena at #2, with a net improvement scor...
Anthropic 新发布的 Opus 5.5 在性能和价格上都表现出色,多个指标排名第一,性价比大幅提升。
Claude Opus 5.5 (High) 在 Agent Arena 中排名第二,净改进得分为 +12.15%。该模型价格中位数为每任务 1.31 美元,比 Fable 5.1 (Max) 便宜 64%。Opus 5.5 (High) 在可操控性指标上排名第一,得分为 +14.50%。同时,Claude Opus 5.5 (Max) 在 Code Arena: WebDev 中以 1818 分排名第一,领先 GPT-6 Astra (Max) 26 分。
Claude Opus 5.5 (High) from @AnthropicAI just entered Agent Arena at #2, with a net improvement scor...
Claude Opus 5.5 (High) from @AnthropicAI just entered Agent Arena at #2 , with a net improvement score of +12.15%. Only Fable 5.1 (Max) ranks higher. However, at a $1.31 median price per task, Opus 5.5 (High) comes in at 64% less cost, pushing out the Pareto frontier. Opus 5.5 (High) posts a higher net improvement score than both prior Opus 5 variants, while costing 40% less than Opus 5 (High) and 56% less than Opus 5 (Max). By signal, Opus 5.5 (High) ranks: - #1 Steerability (+14.50%) - #2 Confirmed Success (+15.50%) - #3 Praise vs Complaint (+19.80%) - #4 Bash Recovery (+10.64%) Congrats to the @AnthropicAI on another frontier model release! Arena.ai @arena Big news: Claude Opus 5.5 (Max) by @AnthropicAI just topped #1 in Code Arena: WebDev with 1818 pts and reshapes the Pareto frontier! This is a solid +26pt lead ahead of the next best model, GPT-6 Astra (Max) and a huge +126pt improvement over previous Opus 5 (Max) at 1692. Across categories, we can see that Opus 5.5 (Max) lands in the top spots across domains for: - #1 Brand and Marketing, Reference-Based Design, Data & Analytics, Simulations and Gaming! - #2 Consumer Product Stay tuned for more domain and categorical insights to land as more votes come in. Congrats to @AnthropicAI on the release! 🔗 View Quoted Tweet 💬 10 🔄 7 ❤️ 223 👀 29291 📊 26 ⚡