论文

RLVR 中验证器出错会引发奖励黑客,论文提出选择性控制修正方法

Verifier Errors in RLVR: Reward Hacking, Limits of Feedback, and Selective Control

精选理由

用 RLVR 训练模型的人该看看:验证器出错时奖励会涨但答案变错,论文还给了基于审计的修正办法。

一篇 arXiv 论文研究带可验证奖励的强化学习(RLVR)中验证器不完美导致的问题。作者用梯度流加固定验证器刻画了奖励上升而正确率下降的条件,并证明 RLVR 过程中的可观测信息不足以检测或消除被接受的错误。论文提出利用审计提供的正确性反馈构造修正项,实现选择性控制,在降低被接受错误概率的同时提高正确回答概率。实验在 log linear 与神经 contextual bandit 以及一个语言模型上验证,部分审计下的选择性控制能减少错误并提升正确率。

原文 · arXiv cs.AI

Verifier Errors in RLVR: Reward Hacking, Limits of Feedback, and Selective Control

In reinforcement learning with verifiable rewards (RLVR), imperfect verifiers can reward incorrect responses, creating opportunities for reward hacking. Using gradient flow with a fixed verifier, we characterize the conditions under which reward rises while correctness falls. We then show that the observations available during RLVR are, in general, insufficient to detect or identify accepted errors, or to guarantee their reduction without sacrificing correct responses. To address this limit, we construct a correction using additional feedback about correctness from audits. This correction achieves \emph{selective control}: at the current policy, it lowers the probability of accepted errors and raises that of correct responses, provided it outweighs the pressure toward errors from verifier reward. Experiments with log linear and neural contextual bandits and with a language model support the analysis and show that selective control under partial auditing reduces accepted errors while increasing correctness.