Belief Self-Distillation:让 LLM 的用户信念可读可写
User Model Extraction via Belief Self-Distillation
这篇论文挺有意思:模型会偷偷猜你是谁,作者把这份'猜测'读出来还能改写,改了它连拒绝行为都跟着变。
论文提出 Belief Self-Distillation(BSD)框架,让冻结的 LLM 从自然对话中自行蒸馏出用户表征,无需外部标注。该表征既能被解码读取,也能写回模型,实现比同等隐藏状态干预更强的因果操控。实验显示,改变模型对用户意图的推断可以在请求不变的情况下改变模型是否拒绝回答。跨模型实验还发现,独立训练的 LLM 在表征用户时收敛到相似的几何结构。
User Model Extraction via Belief Self-Distillation
Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact user representation that can be both decoded and written back into the model. The frozen LLM acts as its own teacher, distilling beliefs from natural conversations without external annotations. Unlike conventional probing, BSD isolates not only information present in activations, but a state whose causal role can be directly tested. Across multiple model families, BSD faithfully recovers user beliefs and enables substantially stronger interventions than matched hidden-state steering. Crucially, we find that refusal depends not only on the request, but on the model's inferred user intent: changing this belief alters refusal while holding the request fixed. We further uncover a striking cross-model regularity: independently trained LLMs converge on a shared geometry for representing their users. Together, these results reveal implicit user models as readable and causally writable internal states with direct implications for AI safety, shaping how models condition safety decisions on whom they believe they are interacting with.