SelfSearch:智能体自改 harness,$4.03 达到 Codex 水平
DeepSeek V4 Flash 加自改 harness,花 4 美元搜索就在 Terminal-Bench 上追平 Codex,SWE-bench 还省了近四成开销,做智能体的都该看看。
SelfSearch 让编码智能体利用自身修改历史记录(包含推理、工具操作和结果)不断改写自己的指令、工具和流程。用 DeepSeek V4 Flash 搜索出的自改 harness 在 Terminal-Bench 2.1 上解决 82.0% 的任务,搜索成本仅 $4.03,与同设置下九个 harness 公开对比中的榜首 Codex 持平。在全部六个模型-基准组合中,群体平均成功率均上升,单个智能体最高提升 11.2 分。在 SWE-bench Multilingual 上,一个进化后的智能体提升 5.0 分,且在双方都能解决的任务上花费减少 38.5%。
Learn to optimize your own harness, folks. You can squeeze much more performance and value from a harness. This work claims that a self-modified harness matched Codex for $4.03. Specifically, a coding agent rewrote its own harness until it solved 82.0% of Terminal-Bench 2.1 with DeepSeek V4 Flash, without any task reward during the search. The search cost was $4.03. That matches Codex, the top harness in a public nine-harness comparison run under the same settings. SelfSearch has agents modify their own instructions, tools, and procedures using records of earlier self-modification attempts. Each record holds the reasoning, tool actions, and outcomes. The modified agent then becomes the next improver. Population-mean success rises in all six model-benchmark settings, with single agents gaining up to 11.2 points. On SWE-bench Multilingual, one evolved agent gains 5.0 points and spends 38.5% less on tasks both versions solve. Paper: arxiv.org/abs/2609.37968 Chat with Paper: academy.dair.ai/papers/selfsea… 💬 14 🔄 3 ❤️ 28 👀 2601 📊 19 ⚡