语言模型欺骗研究的因果框架

From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research

精选理由

这篇论文提出了语言模型欺骗研究的因果框架,区分了欺骗行为与欺骗机制,通过实验验证了看似欺骗的行为不一定有相应机制支持。

AI 摘要

研究人员引入了一个因果分类法,区分了先前的承诺与回顾性报告、模型偏好与实际输出、虚假偏好与误导接收者的效用敏感性,以及欺骗行为与产生它的目标或策略的来源。研究团队在两个开源模型家族中测试了这些区分。在受控的猜谜游戏和股票交易实验中,他们发现看似欺骗的行为可以在没有相应提议机制的情况下产生,而其他干预措施提供了直接证据,表明接收者的信息状态可以因果性地影响欺骗偏好。

原文 · arXiv cs.AI

From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research

Research and news coverage of language-model deception increasingly attributes human-like mental-state concepts to language models. Such claims can blur the distinction between behavior that looks deceptive and a mechanism that is actually deceptive. We introduce a causal taxonomy separating prior commitment from retrospective report, model preference from realized output, false preference from sensitivity to the utility of misleading a recipient, and deceptive behavior from the provenance of the objective or strategy producing it. We test these distinctions in two open-weight model families. Across controlled guessing-game and stock-trading experiments, we find that deceptive-looking behavior can arise without the corresponding proposed mechanism, while other interventions provide direct evidence that recipient information state can causally affect deceptive preference. These results show that deceptive behavior can provide evidence for a deceptive mechanism. But even evidence for such a mechanism does not establish model agency in the deception.

语言模型欺骗研究的因果框架 · AITOP · AI热报