模型精选

Agent Arena 实测 Jev Router:4700 场智能体会话表现如何

精选理由

Agent Arena 跑了 4700 多场真实会话测 Jev Router,结论挺直接:比 DeepSeek V4.1 Flash 贵 38%,但根据反馈换模型的能力接近 Claude Opus 5.5。

Agent Arena 用超过 4700 场真实智能体会话评测了 typesafeai 的 Jev Router。在性能接近 DeepSeek V4.1 Flash (Max) 的情况下,Jev Router 成本高出 38%,请求延迟中位数为后者的 1.7 倍,未改善当前 Pareto 前沿。其路由选择多为高效模型,最常选 DeepSeek V4.1 Flash,也频繁选择 GPT-6.1 Sol 和 GPT-6 Luna。最大亮点是可操控性得分 +10%,接近 Claude Opus 5.5 (High) 的 +10.48,能根据用户反馈切换到更强的模型。

原文 · lmarena.ai

We evaluated Jev Router by @typesafeai on Agent Arena. Our tests spanning more than 4,700 real-world agentic sessions show: 1. Jev Router does not improve on the current Pareto frontier. For similar performance as DeepSeek V4.1 Flash (Max), the solution costs 38% more and its median model request latency is 1.7x higher. 2. However, it does mostly route to Pareto efficient models. The most LLM it picks is DeepSeek V4.1 Flash, with GPT-6.1 Sol and GPT-6 Luna also being frequent choices. 3. Jev Router’s key strength is steerability. Its score of +10% almost matches Claude Opus 5.5 (High) (+10.48). This highlights the benefit of Jev effectively routing to stronger LLM in response to user feedback. More insights in the thread below. 💬 9 🔄 0 ❤️ 31 👀 4683 📊 12 ⚡

  • Artificial Analysis10-06 13:45原文