论文

RollVerify:用验证修复机制兼顾长尾 rollout 强化学习的效率与准确率

RollVerify: Bridging Efficiency and Accuracy in Long-Tail Rollout Reinforcement Learning

精选理由

训练 LLM 做 RL 时长尾 rollout 又慢又掉点?这篇 arXiv 论文的 RollVerify 提前验证修复样本,效率效率、准确率都不丢。

论文针对 RL 训练中长尾 rollout 导致的 GPU 空泡问题,提出基于 partial rollout 的轻量框架 RollVerify。框架引入 off-policy 偏移度量 OPS 量化部分生成轨迹的偏差,并通过序列级和 token 级验证截断无效轨迹后缀,在样本进入训练前主动修复。在数学推理和工具辅助数学推理实验中,RollVerify 的准确率与全 on-policy 训练相当,同时降低训练成本。代码生成任务的初步结果提供了数学之外的佐证。

原文 · arXiv cs.AI

RollVerify: Bridging Efficiency and Accuracy in Long-Tail Rollout Reinforcement Learning

Reinforcement learning is crucial for improving large language models' reasoning and generalization. It relies on massive rollouts whose lengths become increasingly long-tailed as context windows grow. In on-policy training, these long-tail rollouts can result in GPU bubbles, reducing system utilization and limiting RL scalability. Asynchronous or partial-rollout methods improve throughput by relaxing synchronization, but inevitably introduce stale off-policy samples (trajectories) that may hurt final accuracy. Existing approaches mainly mitigate this off-policy issue by reweighting off-policy samples during training, yet they can still leave a performance gap compared to fully on-policy training. In this work, rather than passively reweighting samples during training, we propose RollVerify, a lightweight RL framework built on partial rollout that actively verifies and repairs samples before they enter training. Specifically, it introduces an off-policy shift metric OPS, to quantify the off-policy deviation of partially generated trajectories. Guided by the OPS constraint, RollVerify performs both sequence-level and token-level verification to identify and truncate invalid suffixes of trajectories. This yields high-quality samples that protect the models' accuracy while preserving the efficiency gains of partial rollout. Experiments on mathematical and tool-assisted mathematical reasoning show that RollVerify achieves accuracy comparable to on-policy training while reducing training cost. Additional code-generation results provide preliminary evidence beyond mathematics.