模型多源确认

Cognition 发布 SWE-2 模型,在编码基准测试中表现强劲

Cognition's new SWE-2 (a Kimi K3 post-train) coming in with some powerful coding benchmarks. "On F...

精选理由

朋友发了 Cognition 新的 SWE-2 模型,这个模型在编码方面很厉害,和 Fable 5.1 差不多,但成本更低。

Cognition 的 SWE-2 模型(基于 Kimi K3 后训练)在 FrontierCode 基准测试中取得 50.00% 的分数,与 Fable 5.1 相当,但成本降低了 64%。该模型在多项评估中与前沿模型表现相当,成本可降低 70%。

原文 · The Rundown AI

Cognition's new SWE-2 (a Kimi K3 post-train) coming in with some powerful coding benchmarks. "On F...

Cognition's new SWE-2 (a Kimi K3 post-train) coming in with some powerful coding benchmarks. "On FrontierCode, SWE-2 achieves a score of 50.00%... matching Fable 5.1 at 64% lower cost." Cognition @cognition Introducing SWE-2, our closest model yet to the frontier. On leading evals, it scores on par with recent frontier models – at up to 70% lower cost. We scaled RL to multiple trillions of parameters, with a refined recipe that pushes the Pareto curve on both capabilities & cost. 🔗 View Quoted Tweet 💬 3 🔄 0 ❤️ 3 👀 1444 📊 2 ⚡