论文

研究发现模型被指示说谎时推理 token 数量明显增多

Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models

精选理由

3 个推理模型、210 道题的实验:模型被要求说谎时 token 用得更多,数 token 也能当查谎信号,思路挺巧

一项 arXiv 论文让 3 个具备推理能力的模型在 210 道选择题上分别按说真话、说谎话、不顾真假三种 system prompt 作答,题目覆盖分析性、描述性、规范性推理以及道德与非道德领域。结果显示,三个模型在说真话条件下生成的推理 token 都少于说谎和不区分真假两种条件。这意味着当思维链内容不可读、不可信甚至无法访问时,推理 token 数量可以作为区分模型真实与虚假回答的低带宽替代信号。论文同时指出这仍是概念验证,尚未验证对自发欺骗的实例级检测率和对抗压力下的鲁棒性。

原文 · arXiv cs.AI

Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models

Monitoring the chain-of-thought of reasoning artificial intelligence (AI) models remains a key approach to detecting deception and other forms of misbehavior in such models. However, semantic chain-of-thought monitoring depends on reasoning traces being legible and sufficiently faithful to the underlying computations that produced the model's behavior, not to mention accessible. Moreover, there is increasing evidence that chain-of-thought outputs may soon become illegible or unfaithful, if they even remain accessible. Based on cognitive load theory, we investigate a lower-bandwidth signal -- the number of reasoning tokens generated -- which does not require access to the content of the reasoning trace. Three reasoning-capable large language models answered 210 multiple-choice questions -- across analytic, descriptive, and normative reasoning types as well as moral and non-moral domains -- under system prompts instructing them to respond truthfully, falsely, or without regard for truth. Across all three models, truth-directed responding elicited fewer reasoning tokens than both lie-directed and truth-indifferent responding. These findings show that explicitly prompted untruthful response policies can produce robust group-level differences in test-time reasoning-token use. While not yet establishing reasoning-token count as a detector of spontaneous deception or general misalignment, our results are a proof of concept that it can serve as a simple, content-independent candidate signal for differentiating untruthful from truthful model behavior when raw reasoning traces are unavailable or unreliable. Future work should test instance-level detection rates, out-of-distribution generalization, learned deceptive policies, hidden objectives, and robustness under adversarial pressure.