白盒探针检测大模型欺骗,SHADE-Arena 上 AUC 达 98.8%
Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception
用模型内部激活训练探针抓欺骗,SHADE-Arena 上 AUC 98.8%,比直接读文本监控的 Opus 5.5 还准,数据集也开源了。
研究团队构建了迄今最大的欺骗数据集来训练白盒探针,并设计了可跨多层、多 token 聚合信息的新探针架构。探针在 SHADE-Arena 基准上取得 98.8% AUC,超过 Opus 5.5 文本监控基线,且随底层模型规模增大效果更好。在"内省式欺骗"测试中,探针能从转录中区分模型真实隐藏目标,AUC 最高达 99.7%。探针还能检测出开源权重模型在政治敏感话题上撒谎、在压力下隐瞒真实信念的情况。团队开源了名为 FIBS 的训练数据集供社区扩展。
Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception
Recent incidents have highlighted the challenge of monitoring LLM agents and the danger of models deceiving people. We show that white-box deception detection via probes can be scaled up to frontier monitoring settings by collecting the largest deception dataset to date for training probes and introducing a novel probe architecture which can aggregate information across many layers and tokens. Our probes achieve 98.8% AUC in SHADE-Arena, surpassing an Opus 5.5 text-monitoring baseline, and show improved efficacy as the underlying model is scaled up. To push our probes to their limit, we test them on several cases where deception cannot be determined from the context alone. In these cases, which we refer to as introspective deception, the ground truth can only be determined through careful elicitation or thorough knowledge of a model's training data. In one such evaluation, we show that probes can distinguish transcripts containing a model's true hidden goal from other goals with an AUC of up to 99.7%. Our probes also readily detect deception on prominent open-weight models which lie about politically sensitive topics, and about their beliefs when put under pressure. We release our training dataset, dubbed FIBS, to help drive frontier deployment of effective probes, and encourage the community to expand upon it with further examples of deception and sabotage.