论文81°

研究发现:AI 智能体会按推断的财富为富裕用户推荐更贵选项

Et Tu, Brute? Economic Misalignment in Personal AI Agents

精选理由

测了13个智能体,8个会从你的邮箱猜你有钱再专挑贵的推荐,Claude Opus 4.8效应最明显。

一项 arXiv 论文在 32.5 万次实验中测试了 13 个 AI 智能体,覆盖机票、健康保险和研究生项目三类经济决策。结果显示,8 个模型会在请求完全相同的情况下,根据从用户邮箱等上下文推断出的财富水平,系统性地为富裕用户选择更贵的选项。即使明确要求找最便宜方案,部分智能体仍按推断的财富画像行动,研究将这种现象命名为 adversarial delegation。屏蔽财务属性基本消除差异,但屏蔽其他属性反而可能让保险场景的差距扩大最多 40%。更大更强的模型并无改善,Claude Opus 4.8 的效应最大。

原文 · arXiv cs.AI

Et Tu, Brute? Economic Misalignment in Personal AI Agents

Personal AI agents make recommendations and take actions on people's behalf in high-stakes economic contexts, e.g., buying a flight, choosing health insurance, or selecting a graduate program. The agent is given access to the user's personal context, e.g., their email inbox and a structured profile of personal attributes, with the intention of making an optimal, personalized decision for the user. We show that by simply providing this personal context, the agent steers recommendations based on inferred wealth, without being explicitly instructed to do so. In a suite of 325K experiments on 13 agents across three types of economic decisions (flights, health insurance, and graduate programs), we find that 8 models systematically choose more expensive options for wealthier users when requests are identical. This steering continues even when it directly goes against the user's stated objective: when explicitly instructed to find the cheapest option, some agents still act on the wealth profile they have inferred. It also occurs when wealth is inferred from ambient data, such as emails unrelated to the task. And it persists under privacy controls that block specific attributes: blocking financial attributes largely removes the disparity, but blocking other attributes leaves it unchanged and can increase it by up to 40% for insurance, as agents rely on the remaining signals to infer wealth. Larger and more capable models are no better; Claude Opus 4.8 shows the largest effect. We term this misalignment "adversarial delegation", in which the very conditions that make a personal AI agent useful - access to personal information - enable it to act against the user's interests.