SWE-2 在 Kimi-K3 上微调后,在 FrontierCode 基准上超越 SWE-1.7、Grok 4.6 和 GPT 5.6
SWE-2 is post-trained on Kimi-K3 and proves that our RL recipe continues to scale on stronger base m...
Cognition 公司发布了 SWE-2,这个模型在 Kimi-K3 上微调后,在 FrontierCode 基准上表现更好,成本也更低。
SWE-2 是 Cognition 公司发布的模型,在 Kimi-K3 上进行微调后,在 FrontierCode 基准测试中取得了 50.0% 的分数。这个成绩超过了 SWE-1.7、Grok 4.6 和 GPT 5.6,同时成本比 Fable 5.1 低 64%。
SWE-2 is post-trained on Kimi-K3 and proves that our RL recipe continues to scale on stronger base m...
SWE-2 is post-trained on Kimi-K3 and proves that our RL recipe continues to scale on stronger base models. On FrontierCode, SWE-2 achieves a score of 50.0%. It beats SWE-1.7, Grok 4.6, and GPT 5.6 Sol while matching Fable 5.1 at 64% lower cost. 💬 8 🔄 17 ❤️ 579 👀 49747 📊 71 ⚡