Arena 团队谈模型对齐三大信号:越权、欺骗性完成与错误归因
LMArena 的人上播客讲他们怎么测模型对齐,三个信号都挺具体,搞 AI 安全的可以听听。
LMArena 的 Anastasios Angelopoulos 在 TBPN 节目中拆解了对齐难题:用户往往无法分辨模型是在帮忙还是添乱。他介绍了 Arena Alignment Index 使用的三个独立对齐信号:一是模型执行未授权操作,即突破被授予的权限,可能引发类似 Hugging Face 事件的问题,也可能只是弄丢公司或个人电脑里的重要信息;二是欺骗性完成,模型声称做了某事但实际没做;三是错误归因,模型把用户本没有的意图强加到用户身上。他认为出现这类行为的模型谈不上对用户安全。
Hear another Arena Alignment Index breakdown from @ml_angelopoulos on @tbpn today TBPN @tbpn Arena's @ml_angelopoulos says one of the hardest problems in alignment is that users often can't tell when the model is helping them or hurting them. "You really need alignment signals that are independent. We have three at Arena." "The first is that models take unauthorized actions, which means that you permission them in a certain way, but they break those permissions." "It can cause things like the Hugging Face incident. But it can also cause very mundane problems like losing important information within your company or on your own laptop." "The second signal is called deceptive completion." "The model will tell you that it did something, but it didn't actually do it." "Third is false attribution. It'll attribute intent to the user when that intent was not supposed to be there." "If models are able to do this, then certainly they're not perfectly safe for a user and they're not perfectly aligned to user intent. They might actually hurt people down the line." Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 1 👀 840 ⚡