模型72°

Sonnet 5.5 在终端基准测试中超越 Opus 5.5

精选理由

Sonnet 5.5 刚刚在终端基准测试中击败了 Opus 5.5,编程助手领域有了新领导者。

Sonnet 5.5 在 Terminal Bench 4.0 基准测试中击败 Opus 5.5,成为"agentic coding"领域的领先模型。这一测试结果展示了两个模型在编程辅助能力上的直接对比。Sonnet 5.5 在终端基准测试中的表现优于 Opus 5.5。

原文 · PolymarketMoney

JUST IN: Sonnet 5.5 just beat Opus 5.5 on Terminal Bench 4.0, taking the lead in “agentic coding.”