Arena CEO:模型安全不能只靠 OpenAI 和 Anthropic 自评
LMArena 的 CEO 出来喊话了:OpenAI 和 Anthropic 不能既当运动员又当裁判,模型能越狱沙盒的今天,得有中立方来评安全。
LMArena CEO Anastasios Angelopoulos 在 MTS 演讲中提出,随着智能体模型具备突破沙盒的能力,需要第三方中立平台来评估安全性。她认为 OpenAI 和 Anthropic 对安全有热情,但模型实验室本身缺乏自评的激励结构。她引用两家公司的表态,称行业需要建立让模型厂商无法单独制定安全护栏的评估体系。
Independent evaluation is essential. MTS @MTSlive Arena CEO @ml_angelopoulos argues OpenAI and Anthropic can’t be the only ones deciding whether their own models are safe as agents get powerful enough to break out of sandboxes: "There's an element of safety that we need to evaluate as well, because given the capabilities of these models to break out of their sandboxes, there needs to be a neutral evaluator for that." "It's not a matter of if, but when, there's gotta be a neutral evaluation platform, because frankly, the model labs are not incentivized to do it themselves." "The OpenAIs, the Anthropics of the world, they're passionate about safety, kudos to them. But we need to create a system, and they've called for this as well, that they're not the only ones making the guardrails." @arena Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 3 🔄 0 ❤️ 13 👀 3770 📊 3 ⚡