模型多源确认83°

Mistral Large 4 超越 Opus 5.5 和 GPT-6 Astra

精选理由

Mistral Large 4 在网络安全基准上击败了 Opus 5.5 和 GPT-6 Astra,安全过滤策略更宽松。

Mistral Large 4 在网络安全基准测试中表现优于 Opus 5.5 和 GPT-6 Astra。该模型拒绝了更少的安全任务,而其他两个模型约有 40% 的任务被其安全过滤器阻止。这一性能差异主要源于 Mistral Large 4 在安全过滤策略上的不同。

原文 · Cline

Mistral Large 4 beats Opus 5.5 and GPT-6 Astra on cybersecurity benchmarks, largely because it refuses far fewer security tasks.

Meanwhile Opus and Astra had ~40% of tasks blocked by their own safety filters. https://t.co/gNHAZBV5Ms

  • Simon Willison’s Weblog10-06 18:20原文
  • lmarena.ai06:18原文
  • Mistral AI10-06 13:06原文
  • Guillaume Lample (Mistral)10-06 13:24原文
  • TestingCatalog10-06 13:28原文
  • Arthur Mensch10-06 13:29原文
  • Julien Chaumond10-06 13:30原文
  • Artificial Analysis10-06 13:45原文
  • IT之家10-06 13:56原文
  • OpenRouter10-06 14:02原文