Anthropic Opus 5.5在Agent Arena排名第二
Anthropic新发布的Opus 5.5性能提升显著,成本却大幅降低,重新定义了性能与成本的平衡点。
Anthropic Opus 5.5 (High)在Agent Arena基准测试中排名第二,净改进得分为+12.15%。该模型在成本方面表现出色,每任务中位价格为1.31美元,比Opus 5 (High)低40%,比Opus 5 (Max)低56%。Opus 5.5在可操控性指标上排名第一,达到+14.50%。
Opus 5.5 (High) ranks #2 in Agent Arena and reshapes the Pareto frontier. Opus 5.5 (high) not only improved upon both Opus 5 variants with a higher net improvement score than either, but does so at at 40–56% lower cost: - Opus 5.5 (High): $1.31 median per task / +12.15% net improvement - Opus 5 (Max): $2.98 median per task / +9.58% net improvement - Opus 5 (High): $2.17 median per task / +9.47% net improvement Congrats again to @AnthropicAI ! Arena.ai @arena Claude Opus 5.5 (High) from @AnthropicAI just entered Agent Arena at #2 , with a net improvement score of +12.15%. Only Fable 5.1 (Max) ranks higher. However, at a $1.31 median price per task, Opus 5.5 (High) comes in at 64% less cost, pushing out the Pareto frontier. Opus 5.5 (High) posts a higher net improvement score than both prior Opus 5 variants, while costing 40% less than Opus 5 (High) and 56% less than Opus 5 (Max). By signal, Opus 5.5 (High) ranks: - #1 Steerability (+14.50%) - #2 Confirmed Success (+15.50%) - #3 Praise vs Complaint (+19.80%) - #4 Bash Recovery (+10.64%) Congrats to the @AnthropicAI on another frontier model release! 🔗 View Quoted Tweet 💬 13 🔄 16 ❤️ 190 👀 16515 📊 28 ⚡