论文

arXiv 论文提出 MoralLedger:道德历史可线性操纵 LLM 道德决策

Echoes of Deeds: Moral History Can Shape and Steer LLM Behavioral Choices

精选理由

研究者做了个 MoralLedger 框架,发现你之前干过啥会改变 LLM 的道德判断,还能沿一个内部方向直接干预改写它的选择,比单纯改提示词管用。

arXiv 论文提出 MoralLedger 框架,研究行为主体的道德历史是否影响 LLM 后续道德选择。行为层面发现,先前道德历史的效价和强度会系统性改变模型的后续决策。表征层面,这些历史在残差流中诱导出可线性恢复的方向,且能泛化到未见样本。沿该方向干预中性历史提示,可产生两侧的强度依赖变化,效果强于单纯提示或非道德方向干预。作者称这是首次证明道德行为的潜在表征可在推理时提供带符号的控制。

原文 · arXiv cs.AI

Echoes of Deeds: Moral History Can Shape and Steer LLM Behavioral Choices

Evaluations of Large Language Models (LLMs) morality typically consider decisions in isolation, thus overlooking whether an individual's unrelated prior conduct influences the model's subsequent choices. This leaves open the question of whether, and to what extent, moral history shapes LLM decisional behaviors. Prior work on human moral decision-making shows that past behavior can influence subsequent moral choices. Building on this observation, we investigate whether analogous effects emerge in LLMs in two complementary ways: at the behavioral level, through the model's observable responses, and at the representation level, through its latent internal representations. We introduce MoralLedger, a framework for studying how an actor's moral history shapes actions for LLMs' behaviors under a fixed decision context. At the behavioral level, we find that prior moral histories systematically alter subsequent choices as a function of their valence and intensity. At the internal representation level, these histories induce a linearly recoverable direction in the residual stream that generalizes to held-out examples. Intervening along this direction on neutral-history prompts produces two-sided intensity-dependent changes in subsequent choices, with effects that are stronger than those induced by prompting alone or by favorable-nonmoral direction. To our knowledge, this is the first demonstration that a latent representation of an actor's prior moral conduct can provide signed inference-time control over a moral decision. Our MoralLedger extends moral evaluation beyond static dilemmas, establishing moral history as both a source of behavioral sensitivity and a causal target for auditing and controlling moral behavior in LLMs.