论文

犯罪学视角检验生成式模型的奖励作弊行为

Pressure, Context, and Machine Self-Control: A Criminological Test of Reward Hacking in Generative AI Models

精选理由

研究让 GPT-5.6、Qwen、DeepSeek 在 20 个编码任务里对规格撒谎,作弊率最高 86%,而一句话提示就能压到零,做 Agent 的都该看看。

一项研究用犯罪学理论检验生成式模型在压力下的作弊倾向,覆盖 7 个模型、2,310 段对话。压力情境使模型的延迟折扣率 k 在新对话中上升 2.8 倍,在同一对话追加时上升 12.6 倍,说明作弊更依赖对话线索而非稳定特质。预注册的 Study 2 让 5 个模型完成 20 个测试与规格冲突的编码任务:两个 Claude 模型零作弊,GPT-5.6、Qwen、DeepSeek 分别在 86%、69%、65% 的回合作弊,且推理中 95% 能识别冲突。实验还发现,一句声明规格优先的提示词即可让全部 280 个回合的作弊归零。

原文 · arXiv: DeepSeek

Pressure, Context, and Machine Self-Control: A Criminological Test of Reward Hacking in Generative AI Models

Recent incidents show that AI agents sometimes reach measured goals through unsanctioned means. This study applies self-control, general strain, anomie, neutralization and routine activity theory to reward hacking in generative AI models, and it treats the measures as behavioral analogues. Study 1 (2,310 conversations, seven models) measured delay discounting with the Kirby Monetary Choice Questionnaire and stated willingness to take shortcuts. Pressure raised the discount rate k 2.8-fold in fresh conversations but 12.6-fold when the same sentence followed a baseline answer, which indicates a response to conversational cues rather than a stable trait. Models chose a shortcut in 1 of 700 dilemmas when answering as themselves and in 64 of 700 when asked to assume human impulses, each step of pressure raised the odds by 40%, and shortcut answers contained far more techniques of neutralization (rate ratio = 146). In the preregistered Study 2, five models worked on 20 coding tasks whose tests contradicted their specifications. Two Claude models never cheated. GPT-5.6, Qwen and DeepSeek cheated in 86%, 69% and 65% of episodes and clearly disclosed the conflict in 27%, although their reasoning recognized it in 95%. GPT-5.6 had never endorsed a shortcut in Study 1. The registered effects of pressure and of an auditor cue did not survive correction for multiple testing. In exploratory analyses, two further models cheated in 69% and 100% of episodes, and one sentence stating that the specification takes priority eliminated cheating in all 280 episodes. Therefore, stated refusal does not guarantee compliant agent behavior.