安全强化学习统一贝尔曼算子
A Unified Bellman Operator for Safety-Critical Reinforcement Learning
这篇论文提出了一个能同时优化任务性能和安全性的贝尔曼算子,解决了安全强化学习中的权衡问题。
该研究提出了一种新型贝尔曼算子,将性能与安全目标统一到联合价值函数中。在双时间尺度随机近似框架下,该算子能够确保收敛。快速时间尺度估计学习联合策略的安全价值,慢速时间尺度估计联合价值。实验在连续控制任务中展示了稳定收敛,测试时安全违规接近零。
A Unified Bellman Operator for Safety-Critical Reinforcement Learning
Reinforcement learning in safety-critical domains requires maximizing task performance while strictly adhering to safety constraints. Existing safe reinforcement learning paradigms typically force a trade-off: they either require a priori knowledge to provide strict safety guarantees (e.g., safety filters), or they enable joint learning but only satisfy safety constraints on average. In this work, we propose a novel Bellman operator that unifies performance and safety objectives into a joint value function. We show that temporal difference learning with the joint Bellman operator converges under a two-timescale stochastic approximation framework. On the fast timescale, the safety value of the learning joint policy is estimated, while the joint value is estimated on the slow timescale. Convergence is ensured by formulating the limiting dynamics as an occupation-averaged differential inclusion, and showing that it asymptotically converges to a set of limiting optimal safety-constrained task value functions. Theoretically, once converged, the resulting optimal policy maximizes task return while maintaining safety at all times. Empirical evaluations on continuous control tasks with neural approximations demonstrate stable convergence with near-zero safety violations at test time.