Meta 发布 RankEvolve:用编译协议约束自动研究智能体
Meta 这篇论文讲了怎么让 Claude Code 和 Codex 组队跑 ML 实验,互相挑错后准确率从 45.8% 拉到 62.5%,写智能体工作流的人别错过。
Meta 论文提出 RankEvolve,用编译后的协议对 ML 实验的每个研究阶段和检查点做强制校验,防止静默 bug(如评测数据泄漏、梯度断连)污染实验。它把 Claude Code 和 Codex 作为独立节点,互相审查和修复对方的改动。相同预算下,两个产品组合将执行准确率从单产品的最高 45.8% 提升到 62.5%。在 HSTU 推荐模型上迭代 12 轮后,MovieLens-20M 的 NDCG@10 比已发表结果提高 4.48%。
Must-read paper from Meta on reliable auto-research agents.
If you let coding agents run ML experiments, one silent bug like leaked eval data or a disconnected gradient can invalidate hours of training and every iteration built on it.
RankEvolve enforces each research phase and gate through a compiled protocol. It runs Claude Code and Codex as separate nodes that review and repair each other's changes.
At a matched budget, combining the two products raises execution accuracy from 45.8% for the best single product to 62.5%.
Over twelve iterations on the open-source HSTU recommender, it improved NDCG@10 on MovieLens-20M by 4.48% over the published result.
Paper: https://t.co/HzlsOulc5D