Microsoft 提出 Taste-Bench:智能体在任务分叉点仅 59.7% 选对方向
Microsoft 这篇论文做了个 Taste-Bench,专门考智能体在任务分叉时选路的品味,最强模型也就答对 59.7%,还发现加大推理预算没用。
Microsoft 联合多家机构发布 Taste-Bench,测试智能体在长任务中遇到分叉时的选择能力。基准从工程和研究轨迹中的并行尝试与绕路里自动挖掘分叉点,最佳模型答对率仅 59.7%。决定性证据出现在轨迹后段的分叉明显更难,加大推理预算也无法提升准确率。团队将教师模型对结果的判断蒸馏进学生模型,在 held-out 的 SWE-bench Pro 任务上提高了端到端成功率。
Banger paper from Microsoft and colleagues. We talk about human taste being important in this AI era. But for recursive self-improvement, an agent's taste also matters. The big question is: How often do frontier agents pick the better direction at a decision point in a long task? It's apparently just under 60% of the time, according to Taste-Bench. In this benchmark, each question shows a point in a trajectory where several directions are open, and one leads to a better outcome. The forks are mined automatically from parallel attempts and detours in engineering and research runs. The best model answers 59.7% correctly. Forks whose deciding evidence appears later in the trajectory are much harder, and a larger reasoning budget does not raise accuracy. The authors then distill a teacher's judgment of the outcome into a student model, which improves end-to-end success on held-out SWE-bench Pro tasks. Paper: arxiv.org/abs/2609.25804 Chat with Paper: academy.dair.ai/papers/the-tas… 💬 4 🔄 1 ❤️ 9 👀 1503 📊 5 ⚡