论文

研究提出 selective long-horizon refinement,提升终端智能体训练效率与可靠性

Selecting Long-Horizon Trajectories for Reliable and Efficient Terminal-Agent Training

精选理由

训练终端智能体的朋友看这篇:监督窗口选 12K 反而比 16K 好,还能省 30% 训练时间,作者给的选择性精调方法直接可用。

该论文研究终端智能体训练中的监督窗口长度问题,发现在 Terminal-Bench 上 12K token 窗口比 16K 多解题(29±0.7 对 26±0.8),且训练时间少 30%。作者提出 selective long-horizon refinement 方法,先在短前缀上训练,再只用模型认为最可能延续的部分做精调。该方法在 16K 窗口下把成功尝试从 110±2.7 提升到 126±2.1,用一半长窗口数据还能少 23% 训练时间。结果在 Terminal-Bench v2.0(64 到 73)和 OpenThoughts-TBLite(137 到 155)上均有效。

原文 · arXiv cs.AI

Selecting Long-Horizon Trajectories for Reliable and Efficient Terminal-Agent Training

Terminal agents are commonly trained by imitating long teacher trajectories, yet how much of each trajectory to supervise remains unexplored. We study the \emph{supervision horizon}, the number of trajectory tokens retained for training, and show that it is a key design axis for reliability and cost. Reliability improves with longer horizons but saturates: on Terminal-Bench, a 12K-token horizon solves more tasks than 16K ($29\pm0.7$ vs.\ $26\pm0.8$) while requiring 30\% less training time. The horizon also shapes agent behavior: short horizons cause premature termination, intermediate horizons yield productive error recovery, and long horizons induce over-persistence. We analyze this saturation through a bias--complexity bound, in which longer supervision reduces temporal supervision bias but increases finite-sample estimation error from more heterogeneous late-stage histories. Guided by this analysis, we propose \emph{selective long-horizon refinement}, which first trains on short prefixes and then refines only on continuations that are most likely under the warm-start model. It consistently outperforms full long-horizon training. At 16K, it raises successful attempts from $110\pm2.7$ to $126\pm2.1$ and tasks solved in at least six of eight attempts from $9\pm0.7$ to $14\pm0.6$; with half of the long-horizon data, it still reaches $122\pm2.4$ while cutting training time by 23\%. The gains transfer across benchmarks, from $64\pm2.6$ to $73\pm2.1$ on Terminal-Bench v2.0 and from $137\pm2.7$ to $155\pm2.2$ on OpenThoughts-TBLite. For long-horizon supervision, selecting the right trajectories matters more than training on all of them.