模型

欧洲模型在 CyberGym-E2E-AA 基准以 82% 超越 MiMo-V2.6-Pro 和 GPT-6 Luna

精选理由

欧洲模型在 CyberGym 网络安全基准拿了 82%,压过 MiMo-V2.6-Pro 和 GPT-6 Luna,看来欧洲还没掉队

该模型在 CyberGym-E2E-AA 基准上取得 82% 的成绩,领先 MiMo-V2.6-Pro 的 79% 和 GPT-6 Luna 最高的 78%。这是欧洲模型在网络安全智能体评测中首次排到前列。CyberGym-E2E-AA 考察端到端的攻防实操能力,此前该榜单长期由中美模型占据头部。

原文 · Teortaxes

> Its strongest result is on CyberGym-E2E-AA, where it scores 82%, ahead of MiMo-V2.6-Pro (79%) and GPT-6 Luna (max, 78%)

Maybe Europe is not entirely cooked yet

  • DeepLearning.AI10-06 20:34原文