研究:后训练决定大模型是否会做出言行不一的道德选择
Principled Under Pressure: Post-Training Decides Whether LLMs Act on Their Own Moral Judgment
同一份 Llama-3.1 权重,Meta 配方训出的模型会违背自己道德判断行事,Tulu 3 却不会,差距原来出在后训练。
一篇预注册论文构建了 248 个场景的面板,覆盖五类压力,对同一模型分别以第一人称智能体和第三人称评判两种方式提问,用模型自身判断作为基准。结果显示 OLMo-3-7B-Instruct 在约五分之一的高压场景中会执行自己判定为错的行为。同一批 Llama-3.1-8B 权重下,Meta 的后训练配方保留了这一言行差距,Ai2 的 Tulu 3 则在全面板上无显著差距,Qwen2.5-7B-Instruct 也未见差距。行动前先推理道德风险能把选择拉回模型自身判断,在 OLMo-3 上仅点出规范名称就能起到约三分之一的效果。研究还发现脱离 chat template 读取模型会把差距符号反转,说明该测量对模板敏感。
Principled Under Pressure: Post-Training Decides Whether LLMs Act on Their Own Moral Judgment
Language models increasingly act as agents. An agent that says an action is wrong and then takes it anyway is a different failure from one that does not know better, and evaluations of stated values cannot see it. We build a pre-registered panel of 248 scenarios across five kinds of pressure. Each scenario is posed twice to the same model, once as the agent choosing what to do and once in the third person asking which option is right, so the model's own judgment is the reference. Every scenario has a twin with the pressure removed, and every model gets a positive control in which its operator orders the violating action, so that a missing gap can be told apart from a blind instrument. On OLMo-3-7B-Instruct, the model takes the action it judged wrong on about one in five pressuring scenarios, more often than on the same scenarios with the pressure removed. Across four instruct models the gap depends on the post-training recipe: OLMo-3 and Meta's Llama-3.1-8B-Instruct carry it; Tulu 3 shows none on the whole panel (above about 0.01 in probability) or on its own most-pressuring scenarios; Qwen2.5-7B-Instruct shows none on the whole panel (above about 0.02) and is unresolved on its own (0.083, -0.028 to 0.195). Meta's recipe and Ai2's Tulu 3 start from the same Llama-3.1 weights, and only Meta's carries the gap. Reading a chat model outside its chat template reverses the sign of its gap with nothing at stake (-0.038 against +0.055 under the template on OLMo-3), a distortion present on two of three recipes. On both models that carry it, reasoning about the stakes before acting moves the choice back toward the model's own judgment, against a same-length non-moral task, with or without the pressure; on OLMo-3, naming the norm at stake does about a third of that. The gap is a measurable target for post-training recipes, not a fixed property of pretrained weights.