论文78°

谷歌研究发布用于长期任务的agent harness,在定理证明任务上取得新成果

Banger paper from Google on agent harnesses for long-horizon tasks. (bookmark it) Google Research ...

精选理由

谷歌研究团队发布了新论文,用他们的agent harness在数学证明任务上取得了新进展,和之前的模型相比,这个方法在解决复杂问题上有不同优势。

谷歌研究团队构建了一个多agent harness用于解决长期数学证明问题,在TCS-Bench基准测试中,使用Gemini 3.1 Pro和3.7 Flash模型达到71.0%的准确率,并解决了218个Codeforces编程问题。

原文 · elvis

Banger paper from Google on agent harnesses for long-horizon tasks. (bookmark it) Google Research ...

Banger paper from Google on agent harnesses for long-horizon tasks. (bookmark it) Google Research built a many-agent harness for long mathematical proofs, and it produced new results on open problems from FOCS and JMLR papers. Stellar Colosseum works in stages. It explores several proof strategies, waits for a readiness gate before breaking a route into section-level subproblems, and sends each verifier finding back to the section it affects. Inside each stage, candidates are generated in parallel, attacked with targeted falsification, and merged together with their critiques. With Gemini 3.1 Pro and Gemini 3.7 Flash it reaches 71.0% on TCS-Bench, a set of research-level theorem-proving tasks from FOCS, STOC and SODA papers. With execution feedback it solves 218 of 222 Codeforces problems. Paper: arxiv.org/abs/2609.15983 Chat with Paper: academy.dair.ai/papers/stellar… 💬 5 🔄 3 ❤️ 19 👀 1817 📊 10 ⚡