QF3 算法用过滤 Q 梯度加速流策略强化学习训练
QF3: Fast Flow RL with Filtered Q-Gradients
机器人强化学习新算法 QF3,比 FPO++ 训练快 10 倍,人形机器人策略第一次能从零训完直接上真机,做机器人方向的朋友可以看看。
QF3 是一种在线 off-policy 强化学习算法,通过 flow matching 结合 critic 的动作梯度,经由流输出的一步预测反向传播来训练流策略。为避免不可靠更新,它只对接近回放动作的动作维度应用 critic 梯度。QF3 首次实现从零训练人形机器人运动策略并零样本迁移到真机,配合高吞吐训练方案,比近期的 on-policy 方法 FPO++ 获得了 10 倍的墙钟时间加速。该算法还在 ABC-Sim 和 Robomimic 任务上微调了预训练的流操作策略。
QF3: Fast Flow RL with Filtered Q-Gradients
Flow policies have become a standard policy class for learning robot behaviors from demonstrations, but reinforcement learning is still critical for improving pre-trained flow policies or learning them from scratch through interaction. We introduce QF3 (Fast Flow RL with Filtered Q-Gradients), an online off-policy RL algorithm that trains a flow policy with flow matching plus the critic's action gradient, backpropagated through a one-step prediction of the flow's output. To keep updates where the critic and this prediction are reliable, QF3 applies the critic gradient only to action dimensions that stay near the replay action. To our knowledge, QF3 is the first off-policy flow RL method to train humanoid locomotion policies from scratch and transfer them zero-shot to hardware. Paired with a high-throughput off-policy training recipe, it trains humanoid locomotion and motion-tracking policies with a 10x wall-clock speedup over FPO++, a recent on-policy flow RL method. We further apply QF3 to fine-tune pretrained flow-based manipulation policies on both ABC-Sim and Robomimic tasks. These results suggest that QF3 can both learn robot policies from scratch and refine those acquired from demonstrations. Website: https://qf3-rl.github.io/