论文多源确认

MCP 错误信息为人类开发者编写,反而最伤强智能体

MCP Error Messages Written for Developers Hurt the Most Capable Agents Most

精选理由

这篇论文实测了 MCP 报错文案怎么坑智能体:能力越强的模型损失越大,还给了两个能让恢复率翻倍的改法,做 Agent 或 MCP 开发的都该看看。

研究分析了 150 个常用 MCP server 中的 3,001 条错误信息,其中 949 条指示调用者执行下一步操作,一半步骤依赖 server 无法感知的调用方环境。用五个 OpenAI 模型在 Berkeley Function Calling Leaderboard 任务上测试,发现凭证过期时错误信息里的终端命令导致任务恢复率只剩 45%,且损失从 GPT-5.5 的 18 分扩大到 GPT-6 Astra 的 69 分;GitHub 的限流提示 "Wait before retrying." 让恢复率降到 6%。补救方面,MCP 开发者在错误信息中直接指名 server 工具可将恢复率提升到 84%-88%,智能体开发者用一句提示词在模型读取前删除该步骤也能恢复到 82%。

原文 · arXiv: OpenAI

MCP Error Messages Written for Developers Hurt the Most Capable Agents Most

Many Model Context Protocol (MCP) servers wrap web APIs built for human developers, and their error messages tell the reader to run a command, edit a configuration, open a web page or wait. Many agents that read them can only call the server's tools. In 150 widely used MCP servers, 949 of 3,001 error messages tell the caller what to do next, and half of these steps depend on something the server cannot see about the caller. On credential errors, 62 of 67 steps ask for a terminal command, a configuration change or a web page; on rate limits, 20 of 30 say to wait and retry without naming the call to repeat. We tested five OpenAI models that act only through the tools of Berkeley Function Calling Leaderboard tasks, and the agents did what the step said. On expired credentials, a terminal command in the step left 45% of tasks recovered, and the loss it caused grew from 18 points for GPT-5.5 to 69 for GPT-6 Astra. On a rate limit, GitHub's "Wait before retrying." left 6%. We tested two remedies. For MCP developers, naming a server tool in the step raised recovery on expired credentials to 84%, with the login tool in place of the command, and on a rate limit to 88%, with the call to repeat in place of the bare wait. For agent developers, deleting the step with a one-sentence prompt before the model reads it raised recovery on expired credentials to 82%.