SelfSearch:无奖励搜索提升智能体自我改进
SelfSearch: Reward-Free Search for Self-Improving Agents
SelfSearch让智能体用历史经验自我改进,省钱又高效,Terminal-Bench上提升11.2%,SWE-bench降成本38.5%。
研究人员提出SelfSearch方法,让智能体利用历史自我改进记录进行优化,无需下游奖励信号。在六种模型-基准测试中,该方法使种群平均成功率提升,单个智能体在Terminal-Bench 2.1上提高11.2个百分点。在SWE-bench Multilingual上,成功率提升5.0个百分点,同时降低38.5%执行成本。仅需4.03美元搜索成本,DeepSeek V4 Flash就能解决82.0%的Terminal-Bench 2.1任务。
SelfSearch: Reward-Free Search for Self-Improving Agents
Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures. Existing approaches use this ability to search for improved agents through repeated downstream evaluation, which incurs substantial costs and ties the search to the evaluated tasks. We introduce \textbf{SelfSearch}, a reward-free search procedure in which agents modify themselves using records of previous self-improvement episodes. These records capture the reasoning, tool actions, and outcomes of earlier modification attempts, providing concrete experience for improving both task solving and self-modification. Without downstream reward signals during search, SelfSearch improves population-mean success over the initial agent in all six model--benchmark settings, with individual agents gaining up to 11.2 percentage points on Terminal-Bench 2.1. On SWE-bench Multilingual, an agent improves success by \textbf{5.0} percentage points while reducing execution cost by \textbf{38.5}\% on tasks solved by both the initial and evolved agents. SelfSearch achieves competitive task success with evaluation-guided search baselines at lower search cost. With only \textbf{\$4.03} in search cost, it produces a harness that solves \textbf{82.0}\% of Terminal-Bench 2.1 tasks with DeepSeek V4 Flash under the settings of a public nine-harness comparison, matching the top-scoring harness, Codex. These results suggest that experience gained through self-modification can improve agents' downstream capabilities and efficiency.
- ARC Prize09-29 18:51原文