论文官方一手

苹果研究新方法改变大模型行为

How Value Induction Reshapes LLM Behaviour

精选理由

苹果搞了个新方法,能让大模型在对话里更会关心人、更诚实,和普通大模型不一样。

苹果提出价值诱导方法,让大模型在对话中表现出好奇心、开放性等特质。这种方法通过在训练数据中嵌入特定价值观(如乐于助人、无害性),影响模型生成内容时的行为模式。研究指出,诱导某些价值可能无意中改变其他行为,甚至导致模型更易受影响或奉承。

图片来源 · Apple ML Research
原文 · Apple ML Research

How Value Induction Reshapes LLM Behaviour

Conversational Large Language Models are post-trained on language that expresses specific behavioural traits, such as curiosity, open-mindedness, and empathy, and values, such as helpfulness, harmlessness, and honesty. This is done to increase utility, ensure safety, and improve the experience of the people interacting with the model. However, values are complex and inter-related – inducing one could modify behaviour on another. Further, inducing certain values can make models more addictive or sycophantic through language used in the generations, with a potential detrimental effect on the…