模型多源确认83°

Anthropic Opus 5.5在Agent Arena排名第2

ICYMI @AnthropicAI’s Opus 5.5 (High) ranks #2 in Agent Arena and reshapes the Pareto frontier. Op...

精选理由

Anthropic新发布的Opus 5.5性能提升显著,成本却大幅降低,性价比极高。

Anthropic的Opus 5.5 (High)在Agent Arena基准测试中排名第二,净改进得分为+12.15%。该模型以每任务1.31美元的中位价格,比Opus 5 (High)成本低40%,比Opus 5 (Max)成本低56%。在可操控性指标上排名第一,确认成功率排名第二。

原文 · lmarena.ai

ICYMI @AnthropicAI’s Opus 5.5 (High) ranks #2 in Agent Arena and reshapes the Pareto frontier. Op...

ICYMI @AnthropicAI ’s Opus 5.5 (High) ranks #2 in Agent Arena and reshapes the Pareto frontier. Opus 5.5 (high) not only improved upon both Opus 5 variants with a higher net improvement score than either, but does so at at 40–56% lower cost: Arena.ai @arena Claude Opus 5.5 (High) from @AnthropicAI just entered Agent Arena at #2 , with a net improvement score of +12.15%. Only Fable 5.1 (Max) ranks higher. However, at a $1.31 median price per task, Opus 5.5 (High) comes in at 64% less cost, pushing out the Pareto frontier. Opus 5.5 (High) posts a higher net improvement score than both prior Opus 5 variants, while costing 40% less than Opus 5 (High) and 56% less than Opus 5 (Max). By signal, Opus 5.5 (High) ranks: - #1 Steerability (+14.50%) - #2 Confirmed Success (+15.50%) - #3 Praise vs Complaint (+19.80%) - #4 Bash Recovery (+10.64%) Congrats to the @AnthropicAI on another frontier model release! 🔗 View Quoted Tweet 💬 6 🔄 1 ❤️ 96 👀 10918 📊 12 ⚡