GLM-5.3 网络安全能力接近 Claude Mythos:ExploitBench 得分 12% vs 14%
开源的 GLM-5.3 网络攻防能力快追上 Claude Mythos 了,20 美元 token 就能挖出 Chrome 漏洞,Andrew Ng 还给出了防御方的判断。
Andrew Ng 在本周信件中引用 Anthropic 的评测数据:开源模型 GLM-5.3 在 ExploitBench 漏洞利用任务中解决 12%,闭源的 Claude Mythos 为 14%,差距很小。测试中仅花费 20.40 美元的 token 就复现了 Google Chrome 近期的一个安全漏洞。对于开源模型带来攻击能力扩散的担忧,Andrew Ng 的观点不同:他认为这是一个工程问题,防御方在长期博弈中占据优势。
An open model nearly matches Claude Mythos on cybersecurity capabilities: GLM-5.3 solves 12% and Claude Mythos 14% of ExploitBench exploit tasks, per Anthropic. Just $20.40 in tokens found a recent security exploit in Google Chrome. 💸 In this week’s letter, Andrew Ng has a different take from those who worry about these capabilities in the wrong hands: This is an engineering problem, and defenders hold the long term edge 🛡️ hubs.la/Q04z3MJH0 eF #DeepLearningAI g #Cybersecurity i #OpenWeights hts 💬 1 🔄 0 ❤️ 5 👀 653 📊 1 ⚡