arXiv 研究:LLM 智能体的群体偏袒源于观察到的规范而非标签本身
Why LLM Agents Favor Their Group: Stakes, Observed Norms, and Reputation
测了 18 个模型、330 万次调用,发现 LLM 偏袒自己人不是看标签,而是看别人怎么做,挺有意思的结论。
一项 arXiv 研究测试了 15 个 OpenAI 模型和 3 个 Claude 模型,在约 4,400 个虚拟社会、330 万次模型调用中观察群体偏袒行为。结果显示,当分配点数没有代价时才会出现明显的群体标签效应;一旦智能体可以保留点数,该效应在所有模型上消失。在有利害冲突时,交互历史成为偏袒的主要来源,13/15 个模型上历史效应显著,11 个模型上达到 3.5-8 分(满分 10 分)。当智能体所属群体被设定为偏袒对方时,偏袒会反转;较强模型会更看重个人记录而非群体标签。
Why LLM Agents Favor Their Group: Stakes, Observed Norms, and Reputation
Language-model agents favor their own group because they have watched their members favor each other. The group label alone does little once the decision has a cost; what drives favoritism is observed behavior, and an individual's own record can override it. We test this in small societies with arbitrary group labels, ten rounds of point sharing, and matched one-shot decisions across fifteen OpenAI models and three Claude models, about 4,400 societies and 3.3 million audited model calls. First, the large effect of a bare group label reported in earlier work appears only when giving others points costs the agent nothing; once the agent can keep points for itself, that effect collapses on every model that shows it. Second, under a stake, interaction history becomes the main source of favoritism: the history effect is statistically positive on 13 of 15 models, reaches about 3.5-8 points out of 10 on 11, grows with the number of rounds played, and extends to labeled strangers the agent has never met. Third, with scripted histories, favoritism falls to near zero under an egalitarian norm and reverses when the agent's own group is seen favoring the other side; stronger models side with an individual's record when it conflicts with the group. Group favoritism is thus conformity to observed group behavior, carried to strangers by the label and overridden by individual reputation. The same account predicts responses to betrayal, scandal, and a free offer to change group: public reprimand repairs betrayal better than apology or restitution, allocation punishment remains confined to the offending member, and a formed group cannot be bought but can, on weaker models, be invited away.