Claude Opus 5.5 上线 Agent Arena,可投票实测智能体任务
Claude Opus 5.5 by @AnthropicAI is now in the Agent Arena! Your votes drive the @arena leaderboards...
Claude Opus 5.5 进 Agent Arena 了,你的 toughest prompt 能直接给模型出难题投票,还能看到它在 WebDev、Vision 等赛道的对战成绩。
Anthropic 的 Claude Opus 5.5 现已加入 Agent Arena 排行榜,该榜单基于数百万条真实用户提交的长程智能体任务,模型可调用网页搜索、文件系统和终端工具完成复杂工作流。排行榜采用因果追踪方法,以相对平均模型的结果表现来计分。除了 Agent Arena,Opus 5.5 还进入了 WebDev、Text、Vision 和 Document 四个赛道的 Battle Mode。官方称其在多数任务上达到 Claude Fable 5.1 水平,运行成本比 Opus 5 低 40%。
Claude Opus 5.5 by @AnthropicAI is now in the Agent Arena! Your votes drive the @arena leaderboards...
Claude Opus 5.5 by @AnthropicAI is now in the Agent Arena! Your votes drive the @arena leaderboards, head over and bring your toughest prompts. In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. In addition to Agent Arena, @claudeai Opus 5.5 is in Battle Mode for: WebDev, Text, Vision, and Document. Claude @claudeai Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5. Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 6 🔄 0 ❤️ 63 👀 6607 📊 9 ⚡