小米发布MiMo-V2.6,研究多任务代理RL的扩展性
Impressive level of openness on such a large run
小米搞了个叫MiMo-V2.6的模型,在研究多任务代理RL怎么扩展,计算量很大,每步20亿令牌,用了很多提示和滚动,还混合了多种环境,挺有意思的。
小米的MiMo-V2.6正在进行大规模强化学习训练,计算量达到每步约20亿令牌,使用1568个提示和16个滚动,完全异步。训练涉及多任务代理RL和混合环境,并使用基于测试用例和评分标准的奖励进行评估。团队计划在未来几周逐步开源细节。
Impressive level of openness on such a large run
Impressive level of openness on such a large run Fuli Luo @_LuoFuli Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run: mimo.xiaomi.com/rl/ 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 9 👀 1455 📊 1 ⚡