技巧多源确认精选72°

Jev 路由实测对比 GPT-6 Astra:公开基准数字参考价值有限

An interesting discussion is happening in the comments. I think it's worth highlighting. I ran thi...

精选理由

有人拿自建 harness 实测了 Jev 路由,性能能追平 GPT-6 Astra low,代价是跑得更久。文章还讨论了为什么公开基准分数别太当真。

独立研究者用自建 harness 测试 Jev Router 在智能体任务中的表现,成绩与 GPT-6 Astra low 相当,但运行时间更长。Theo 用另一套 scoped tests 测了 DeepSWE,两套测试内容不同,结论有分歧也有共识。作者强调 Jev 是 System One 模型,不应与 GPT-6 Astra 这类 System Two 模型直接对比,更值得探索的是两者结合的路由方案。

原文 · elvis

An interesting discussion is happening in the comments. I think it's worth highlighting. I ran thi...

An interesting discussion is happening in the comments. I think it's worth highlighting. I ran this little experiment because I wanted to measure something I use agents for in the real world. The numbers are promising, but I must acknowledge they may not generalize as well to every task. That needs to be said. It works for my custom harness, but it may not generalize to all cases. A lesson here is that public benchmark numbers don't mean much. What may work for you may not work for me, and vice versa. Theo brought up his scoped tests that he did with DeepSWE. You can see the Jev Router performance is comparable to GPT-6 Astra low, but at the expense of a longer run. Mine and his are very different tests measuring different things. And with some disagreements and agreements. But there are so many ways to test and compare this. I also think the reasoning effort is an interesting use case for routing. Jev, as I have been posting, is a System One model and shouldn't be compared with System Two models. Jev doesn't have the capabilities of, say, a GPT-6 Astra, and so will be limited in what you can do with it. I think the more interesting experiment is combining them, hence my post. Even more interesting is potentially a hybrid model, as I shared in a previous post. Routing is largely unsolved, but we all want it. So I don't think we can dismiss ideas based on our little benchmark tests that say nothing about the real world. I am learning that it's getting extremely hard to measure and compare things these days. It's expensive and time-consuming. I really appreciate people like Theo for testing things and sharing his findings, which is what I tried to do too. I hope OpenRouter can share more on how they are measuring this. We also need to keep an open mind and discuss these ideas openly. We shouldn't encourage dismissing ideas so quickly, especially not in the world of harness engineering. I remain positive and believe these discussions will yield better solutions to these problems. I hope to run even more tests with my custom harness and potentially measure more of these things. As an independent researcher, I will continue sharing these ideas openly and hope to spark interesting discussions around harness engineering. Thanks to those who took the time to reply. 💬 0 🔄 0 ❤️ 1 👀 337 📊 1 ⚡