Claude Opus 5.5 登顶 Artificial Analysis 编程智能体指数,但单任务成本更高
Anthropic 的 Opus 5.5 在编程智能体评测拿第一,三项基准都涨了,价格还降了 20%,不过每任务要烧 13 美元 tokens,用前算算账。
Anthropic 的 Claude Opus 5.5 在 Artificial Analysis Coding Agent Index 拿到 66 分,排名第一,比 Opus 5(60 分)高 6 分,比 Claude Fable 5.1(62 分)高 4 分。三项评测全部提升:Terminal-Bench 4.0 从 54.5% 升到 63.1%,DeepSWE v1.1 从 62.5% 升到 68.4%,SWE-Atlas-QnA 从 62.1% 升到 66.4%。价格降到每百万 tokens 输入 $4、输出 $20(Opus 5 为 $5/$25),缓存读取从 $0.50 降到 $0.20。但单任务成本从 $10.79 涨到 $13.04,因为每个任务消耗约 1560 万 tokens,输出 tokens 约为 Opus 5 的 2.4 倍。
Claude Opus 5.5 is the new #1 in the Artificial Analysis Coding Agent Index, with gains across all three evaluations, though at a higher Cost per Task
At max effort in Claude Code, Opus 5.5 scores 66 on the Coding Agent Index, the highest score we have measured. It is up 6 points against Opus 5 (60) and 4 points against Claude Fable 5.1 (62).
Anthropic has cut Opus pricing to $4/$20 per million input/output tokens, from $5/$25 for Opus 5, and cache reads to $0.20 from $0.50. Even with those reductions, Opus 5.5’s Cost per Task is $13.04, above Opus 5’s $10.79, because it uses substantially more tokens.
Key takeaways:
➤ Improves across all three Coding Agent Index evaluations: Terminal-Bench 4.0 rises to 63.1% from 54.5% for Opus 5, DeepSWE v1.1 to 68.4% from 62.5%, and SWE-Atlas-QnA to 66.4% from 62.1%. The largest gain is on Terminal-Bench, at +8.6 percentage points.
➤ The top score comes at the highest Cost per Task: Opus 5.5’s Cost per Task is $13.04, up 21% from Opus 5 at $10.79. It uses about 15.6 million tokens per task against 11.4 million for Opus 5, including about 2.4× as many output tokens.
➤ Extends the Coding Agent Index vs Cost per Task Pareto frontier: No lower-cost model in our comparison matches Opus 5.5's score. It moves the frontier upward at its high-cost end.
Other model details:
➤ Pricing: $4/$20 per million input/output tokens, down 20% from Opus 5. Cache reads cost $0.20 per million, down 60% from $0.50.
➤ Evaluation setup: Claude Code at max effort, measured on DeepSWE v1.1, Terminal-Bench 4.0 and SWE-Atlas-QnA. The Coding Agent Index gives each evaluation equal weight.