论文精选

MIT 论文:告诉 LLM 智能体规则事实比要求多推理更能改善决策

精选理由

MIT 实测 GPT-4o、Claude 这些模型玩拍卖游戏,发现让它们“多想想”反而更差,一句话说明规则就能把错误率从 4.2% 降到 0.2%,写提示词的都该看看。

MIT 一篇论文测试了 GPT-4o、Claude、Gemini 和 Gemma 在拍卖与匹配博弈中的决策表现,发现为人类竞拍者设计的市场机制规则同样适用于 LLM 智能体。将价格上升简化为“留下或退出”的简单选择后,Gemma 的出价差距从低于价值 5.30 美元缩小到 0.30 美元。在匹配环节加一行说明“拒绝只会重新分配”就把错误率从 4.2% 降到 0.2%。四个模型都存在压低出价、保留利润空间的问题,而要求智能体推理对手行为的提示词反而增加了错误。结论是先修正信息呈现格式、直接陈述规则关键事实,再考虑加推理提示。

原文 · rohanpaul_ai

New MIT Paper: LLM agents decide better when the choice is shown in simple steps or the rule's safe move is stated plainly, and worse when told to reason about opponents.

Market-design rules of thumb built for human bidders carry over to LLM agents, so we can borrow them instead of inventing new prompt tricks.

Honest bids and rankings are always the best move in these auctions and matching games. GPT-4o, Claude, Gemini and Gemma still underbid, often to keep a profit margin.

A rising price with a simple stay-or-exit choice moved Gemma from $5.30 below its value to $0.30 below. A 1-line note that rejections only redirect cut matching errors from 4.2% to 0.2%.

Fix the format and state the key fact before adding reasoning prompts, and judge agents by their choices, since their written plans missed these gains.

Agent prompts should state facts about the rules rather than request more thinking, because facts improved choices while thinking prompts often added errors.