Mistral Large 4 在 DeepSWE 基准测试中领先 GLM-5.3
Mistral Large 4 在多个专业领域基准测试中表现优异,金融和法律任务得分亮眼。
Mistral Large 4 在 DeepSWE 基准测试中得分为 62%,超越 GLM-5.3。该模型在 Finch 金融任务基准测试中得分为 67%,为当前开源模型最佳表现。Le Chonk 在 Harvey 的法律智能体基准测试中得分为 15%。
Mistral Large 4 scores 62% on DeepSWE, outperforming GLM-5.3, according to VentureBeat.
Additionally, it scores 67% on Finch (Financial tasks, SOTA open-weight).
Le Chonk also scores 15% on Harvey’s Legal Agent Benchmark (Legal tasks, SOTA open-weight).
We need a tech report now 👀 h/t @AiBattle_