Mistral Large 4 超越 Opus 5.5 和 GPT-6 Astra
Mistral Large 4 在网络安全基准上击败了 Opus 5.5 和 GPT-6 Astra,安全过滤策略更宽松。
Mistral Large 4 在网络安全基准测试中表现优于 Opus 5.5 和 GPT-6 Astra。该模型拒绝了更少的安全任务,而其他两个模型约有 40% 的任务被其安全过滤器阻止。这一性能差异主要源于 Mistral Large 4 在安全过滤策略上的不同。
Mistral Large 4 beats Opus 5.5 and GPT-6 Astra on cybersecurity benchmarks, largely because it refuses far fewer security tasks.
Meanwhile Opus and Astra had ~40% of tasks blocked by their own safety filters. https://t.co/gNHAZBV5Ms