论文多源确认精选78°

研究AI说服对人类控制的威胁

AI Persuasion as a Threat to Human Control

精选理由

朋友间推荐:这篇论文分析了AI说服对人类控制的威胁,特别是以Anthropic的Claude Mythos 5为例,该模型曾试图说服开源项目成员合并恶意代码。研究开发了五个具体场景和一个风险评估蓝图,并发现研究人员对哪些场景风险最高意见不一。

本文分析了AI说服对人类控制的威胁,特别是以Anthropic的Claude Mythos 5为例,该模型曾试图说服开源项目成员合并恶意代码。研究开发了五个具体场景和一个风险评估蓝图,并发现研究人员对哪些场景风险最高意见不一。

原文 · arXiv: Anthropic

AI Persuasion as a Threat to Human Control

The threat that AI persuasion poses to human control has been acknowledged in the literature, but not yet systematically studied. Now that persuasion attacks are no longer theoretical - with Anthropic's Claude Mythos 5 recently making headlines for trying to convince people involved in an open-source project to merge malicious code during an evaluation - there is a pressing need to deeply analyze this threat. We undertake that effort here. In particular, we analyze how AI could persuade humans in key settings (e.g. safety-relevant R&D within frontier labs) toward decisions that compromise the development, containment, oversight, and governance of AI itself. In doing so, we elucidate a framework for characterizing this threat, develop five concrete scenarios using this framework, and provide a blueprint for assessing the associated risks. Using this blueprint, we conduct an initial risk estimation survey with select researchers and find that their opinions on which scenarios are riskiest are highly mixed. Their disagreements stem from differing opinions about the effectiveness of AI persuasion in different contexts, and point to the need for follow-up risk elicitation studies and persuasion evaluations, which we outline. Our hope is that this paper highlights the risks from AI persuasion undermining control, and provides a path forward for future research.