arXiv 论文:可再推导性决定多级智能体流水线故障后的恢复能力
Re-derivability Decides What a Staged Agent Pipeline Recovers After an Upstream Fault
这篇论文用 gsm_hard 上的 120 题实验算清了一笔账:流水线出错后,把原题重新喂给第一个下游阶段就能挽回 0.394,每题只多花 59.8 个 token,后面的阶段白搭。
一篇 arXiv 论文研究了语言模型智能体多级流水线中,上游故障造成的损失由什么决定,答案是可再推导性,即每个阶段能从原始问题重建多少信息。在 4 个开源权重模型上、每臂 120 道 gsm_hard 题目、温度为 0 的设置下,将原始问题重新暴露给下游阶段使故障下的准确率提升 0.233 至 0.392,最大 Holm 校正 p 值为 2.1×10⁻⁶。在 Qwen3-14B 上,第一个重新接地的阶段以每题约 59.8 个 token 的代价买来 +0.394 的匹配保留率,其后的阶段没有额外收益。论文还发现无故障时注册流水线输给单次直接调用 -0.267 至 -0.317,即多级架构本身有代价。
Re-derivability Decides What a Staged Agent Pipeline Recovers After an Upstream Fault
One variable sets what an upstream fault costs a staged pipeline of language-model agents: re-derivability, how much of what a stage needs it can rebuild from the original problem. Grounding an inspector agent in that problem is worth +0.608 [+0.517, +0.700] to +0.358 over a blind one on four open-weight backbones served with thinking disabled, and on the two Qwen backbones the blind inspector changes no item at all. That head-to-head is exploratory. One deterministic fault enters the first stage, and we re-expose the original problem to $k = 0,\dots,3$ of the downstream stages with agents, items, fault and topology held fixed, on 120 gsm_hard items per arm at temperature zero. Accuracy under fault rises on four of four backbones, from +0.233 to +0.392, the largest Holm-adjusted $p$ being $2.1\times10^{-6}$. A registered kill test rules out tokens. Blanking every word holds the word slots fixed, and retention tracks the visible fraction on four of four, climbing from 0.221 to 0.692 on the primary. Those two families are confirmatory and everything else here is exploratory. The interaction excludes zero on two of four backbones under the registered pipeline, four of four under a three-stage pipeline, and three of four under full message history, the primary at +0.317. On Llama-3.1-8B the fault carries no detectable cost at any dose, so the other three carry every claim about what a fault costs. Re-derivability also sets what the architecture costs, and no decomposition we measured reliably beats one direct call. With no fault injected the registered pipeline loses to that call by -0.267, -0.125 and -0.317, and on Phi-4 reads +0.058 at $p = 0.118$, which the test fails to separate from zero. The repair that works is cheap and front-loaded: the first re-grounded stage buys +0.394 of matched retention for +59.8 tokens per item on Qwen3-14B, and the stages after it buy nothing.