欧洲模型在 CyberGym-E2E-AA 基准以 82% 超越 MiMo-V2.6-Pro 和 GPT-6 Luna
精选理由
欧洲模型在 CyberGym 网络安全基准拿了 82%,压过 MiMo-V2.6-Pro 和 GPT-6 Luna,看来欧洲还没掉队
该模型在 CyberGym-E2E-AA 基准上取得 82% 的成绩,领先 MiMo-V2.6-Pro 的 79% 和 GPT-6 Luna 最高的 78%。这是欧洲模型在网络安全智能体评测中首次排到前列。CyberGym-E2E-AA 考察端到端的攻防实操能力,此前该榜单长期由中美模型占据头部。
原文 · Teortaxes
> Its strongest result is on CyberGym-E2E-AA, where it scores 82%, ahead of MiMo-V2.6-Pro (79%) and GPT-6 Luna (max, 78%)
Maybe Europe is not entirely cooked yet
- DeepLearning.AI10-06 20:34原文