模型官方一手精选78°

小米 MiMo 团队直播新模型 MiMo-V2.6 的 RL 训练过程

我操,太整活了! 小米 MiMo 团队居然在直播他们新模型 MiMo-V2.6 的RL 训练过程,看起来是一直在进行自动化的评估和测试。 然后一个特别的点是,它显示了成本,而且成本是实时显示的。 ...

精选理由

小米团队直播新模型 MiMo-V2.6 的 RL 训练过程,实时显示成本,让你看到训练一个模型到底有多费钱。

小米 MiMo 团队直播新模型 MiMo-V2.6 的 RL 训练过程,实时显示成本。训练过程包含自动化的评估和测试,展示了模型训练的成本消耗。

原文 · 歸藏(guizang.ai)

我操,太整活了! 小米 MiMo 团队居然在直播他们新模型 MiMo-V2.6 的RL 训练过程,看起来是一直在进行自动化的评估和测试。 然后一个特别的点是,它显示了成本,而且成本是实时显示的。 ...

我操,太整活了! 小米 MiMo 团队居然在直播他们新模型 MiMo-V2.6 的RL 训练过程,看起来是一直在进行自动化的评估和测试。 然后一个特别的点是,它显示了成本,而且成本是实时显示的。 你可以看到训一个模型到底有多费钱,每一秒都在涨那个钱! Fuli Luo @_LuoFuli Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run: mimo.xiaomi.com/rl/ 🔗 View Quoted Tweet 💬 2 🔄 0 ❤️ 3 👀 1119 📊 2 ⚡