Mistral 发布 1T 参数多模态模型 Mistral Large 4,编程与安全成绩对标 GLM-5.3
Mistral 这次真支棱起来了,1T 参数只激活 49B,用 4000 块芯片就训出能跟 GLM-5.3 掰手腕的模型,欧洲总算有能打的了。
Mistral 发布 Mistral Large 4(代号 Le Chonk),为原生多模态模型,总参数 1T、激活 49B。预览版在 DeepSWE v1.1 得分 61.7%,与 GLM-5.3 的 61% 持平,Kimi K3 仍以约 68% 领先。在 Artificial Analysis Cyber Index 上得分 50,超过 GLM-5.3 的 36;CyberGym-E2E-AA 得 82%,高于 MiMo-V2.6-Pro 的 79%。该模型在欧洲自建数据中心从头训练,仅用 4,000 块 NVIDIA Grace Blackwell Superchips。
This is a surprising great release. Le Chaton fat is real! Mistral Large 4 puts Europe back in contention on coding and cybersecurity.
Didnt expect Mistral to compete with GLM5.3!
“Le Chonk” is a natively multimodal model with 1T parameters and 49B active. The preview already delivers:
- DeepSWE v1.1: 61.7%, roughly level with GLM-5.3 at 61%. Kimi K3 remains ahead at around 68% in Mistral’s comparison. - Artificial Analysis Cyber Index: 50, matching GLM-5.3-Flash and beating GLM-5.3’s 36. - CyberGym-E2E-AA: 82%, ahead of MiMo-V2.6-Pro’s 79%.
Mistral says it trained the model from scratch in its own European datacenters. NVIDIA says only 4,000 Grace Blackwell Superchips powered the training! Which is impressive.
Learning: you can compete with much less compute. Which is crazy. Congrats Mistral!