论文

RoVR-GSPO双通道鲁棒策略优化器发布

Dual-Channel Robust Group-Relative Policy Optimization via Advantage and Sequence-Weight Estimation

精选理由

新发布的RoVR-GSPO优化器解决了奖励和序列权重敏感问题,在多个任务上超越GSPO,对异常值更鲁棒。

RoVR-GSPO是一种双通道鲁棒优化器,解决了相对策略优化中的局部异常值敏感问题。奖励通道结合鲁棒参考估计与有界残差信用,比率通道使用可微分SoftRoVR聚合构建鲁棒序列权重。在数学推理、长文本摘要和工具调用标注任务上,RoVR-GSPO相比GSPO实现了一致的性能提升,同时对奖励污染和token比率异常表现出更强的鲁棒性。

原文 · arXiv cs.AI

Dual-Channel Robust Group-Relative Policy Optimization via Advantage and Sequence-Weight Estimation

Group-relative policy optimization relies on reward-derived advantages and sequence-level likelihood weights, both of which can be sensitive to localized outliers. Extreme rewards can collapse the contrast among clean responses after group normalization, while token-level log-ratio perturbations can alter sequence weights and clipping decisions. We introduce RoVR-GSPO, a dual-channel robust optimizer that addresses these failure modes separately. Its reward channel combines robust reference estimation with bounded residual credit, while its ratio channel uses differentiable SoftRoVR aggregation to construct robust sequence weights. We provide stability and efficiency analyses for both channels. Experiments on mathematical reasoning, long-context summarization, and tool-call annotation show consistent improvements over GSPO, while controlled perturbation studies demonstrate stronger robustness to reward contamination and token-ratio anomalies.