模型多源确认76°

Mistral Large 4 在 DeepSWE 基准测试中领先 GLM-5.3

精选理由

Mistral Large 4 在多个专业领域基准测试中表现优异,金融和法律任务得分亮眼。

Mistral Large 4 在 DeepSWE 基准测试中得分为 62%,超越 GLM-5.3。该模型在 Finch 金融任务基准测试中得分为 67%,为当前开源模型最佳表现。Le Chonk 在 Harvey 的法律智能体基准测试中得分为 15%。

原文 · TestingCatalog

Mistral Large 4 scores 62% on DeepSWE, outperforming GLM-5.3, according to VentureBeat.

Additionally, it scores 67% on Finch (Financial tasks, SOTA open-weight).

Le Chonk also scores 15% on Harvey’s Legal Agent Benchmark (Legal tasks, SOTA open-weight).

We need a tech report now 👀 h/t @AiBattle_

  • OpenRouter10-06 14:02原文
  • kimmonismus10-06 14:18原文
  • Guillaume Lample (Mistral)10-06 13:24原文
  • Julien Chaumond10-06 13:30原文
  • AIGCLINK10-06 14:15原文
  • Thomas Wolf10-06 14:34原文
  • Simon Willison10-06 14:35原文
  • Clement Delangue10-06 17:36原文
  • Mistral AI10-06 13:06原文
  • Arthur Mensch10-06 13:29原文