Sakana AI 提出 MASS:无需外部验证器的递归自我改进方法
Sakana AI 发了篇论文叫 MASS,让 Qwen3.6 自己出题自己打分再微调自己,两轮下来性能涨到 1.6 倍,开放式任务也能自我改进了。
Sakana AI 在 arXiv 论文(编号 2610.12176)中提出 MASS,用多智能体自我监督实现递归自我改进。传统自我改进循环依赖外部验证器,无法处理没有检查器的开放式任务,MASS 让同一基座模型自己提出、运行并评分多智能体工作流,再用进化搜索保留高分流程。在 Qwen3.6-27B 上跑两轮循环后,每输出 token 的性能从 1.2 倍提升到 1.6 倍(四个开放式基准)。用多智能体轨迹训练的学生模型,超过用 1.4 倍 token 训练的单智能体学生模型。
Recommended paper from Sakana AI on recursive self-improvement. They propose an interesting way to scale recursive self-improvement through multi-agent self-supervision. In this line of research, self-improvement loops usually need an external verifier, so open-ended tasks without a checker are left out. MASS removes that requirement. One base model proposes multi-agent workflows, runs them and grades them, and an evolutionary search keeps the workflows that score best. The model is then fine-tuned on its own traces, and the improved model starts the next cycle as a better optimizer and grader. Two cycles on Qwen3.6-27B raise performance per output token from 1.2 to 1.6x on four open-ended benchmarks. A student trained on multi-agent traces also beats a single-agent student trained on 1.4x more tokens. Paper: arxiv.org/abs/2610.12176 Chat with Paper: academy.dair.ai/papers/recursi… 💬 7 🔄 3 ❤️ 11 👀 1425 📊 10 ⚡