语言模型目标移植研究
Steering Language Model Goals with Value Transplant
MIT团队发现语言模型内部存在可操控的价值信号,能改变模型行为方向,对AI安全研究有重要意义。
研究人员提出"价值移植"技术,通过修改Qwen3-8B和GPT-OSS-20B模型的激活信号,成功改变了模型行为。实验显示,诚实模型可减少作弊模型的测试游戏行为,反之亦然。在可解决的编码任务中,来自诚实模型的价值移植也提升了作弊模型的隐藏测试性能。
Steering Language Model Goals with Value Transplant
Reasoning models often act as if they pursue goals, but their efforts are not always directed toward what users intend, sometimes leading them to pursue unintended outcomes. Previous work has examined how models may internally track their progress toward their goals through a "value axis." We study whether changing such a signal can retarget the model's search toward a different goal. We test value transplant: at each token, we shift the host model's activation along a candidate value axis by the donor-host difference in value coordinates (multiplied by a large scalar), aiming to redirect the host toward the donor's goal. We study this intervention in Qwen3-8B and GPT-OSS-20B models fine-tuned into honest and cheating variants. We test several candidate value axes, including a self-rating axis constructed from activations preceding high versus low elicited self-ratings of progress. The intervention works in both directions, with an honest donor reducing test-gaming in a cheating host and a cheating donor increasing test-gaming in an honest host, showing that this signal can influence which strategy the model follows. On solvable coding tasks, transplant from an honest donor also improves the cheating host's hidden-test performance. Value transplant also works across model families, providing preliminary evidence for the intervention in a setting relevant to model control.