论文

FERPO:前熵正则化策略优化算法

FERPO: Forward Entropy-Regularized Policy Optimization

精选理由

新提出的FERPO算法解决了传统强化学习中critic值预测不准确的问题,在连续控制任务中表现优异。

研究人员提出FERPO算法,一种在线强化学习方法,在MuJoCo Playground和ManiSkill基准测试中展现出竞争性性能和样本效率优势。该算法通过前KL目标函数鼓励探索多个高价值模式,避免了传统反向KL目标可能只选择目标分布子集的问题。FERPO使用自归一化重要性采样估计前KL目标,计算速度比REPPO算法更快。

原文 · arXiv cs.LG

FERPO: Forward Entropy-Regularized Policy Optimization

Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives, potentially leading to unreliable policy updates. We propose Forward Entropy-Regularized Policy Optimization (FERPO), an on-policy maximum entropy reinforcement learning algorithm that performs policy improvement using critic values without differentiating the critic with respect to actions. FERPO derives an optimal target action distribution from a policy-improvement objective regularized by entropy and Kullback-Leibler (KL) divergence. We then fit the actor to this target by minimizing a forward-KL objective, estimated using self-normalized importance sampling (SNIS) with actions drawn from the rollout policy. By limiting the target distribution's deviation from the rollout policy, the KL regularization helps keep these importance weights well behaved. In contrast to reverse-KL objectives, which can favor a subset of the target distribution's modes, the forward-KL objective encourages coverage of multiple high-value modes and thereby promotes exploration. Experiments and ablations on MuJoCo Playground and ManiSkill show competitive performance and sample-efficiency gains. Computational benchmarks also demonstrate faster actor updates than Relative Entropy Pathwise Policy Optimization (REPPO).