评论者谈模型能力边界:尚无主动夺权迹象
一位 AI 安全观察者在 X 上聊模型会不会夺权,他拿 Diplomacy 和 Factorio 这些游戏环境当论据,观点挺有意思。
一条 X 帖子讨论模型能力的锯齿状分布,指出目前不存在能训练出“胜过笨蛋”的 RLVR 环境。作者提到 Diplomacy、Stratego、Factorio 这类游戏环境可作为替代测试场景。作者认为,如果模型真的产生了夺取权力的连贯目标,应该已经能看到其至少在考虑此事的迹象。
I am actually uncertain on this point. Yes, models are jagged, there is no RLVR env "outsmart a village idiot". But there is something like Diplomacy and Stratego and Factorio. I think *if* they had a coherent *desire* to take over, we'd see signs of at least considering it. https://t.co/ZeQlO9Qyn5