论文精选

表现强化学习的局部与全局稳定性研究

Local and Global Stability in Performative Reinforcement Learning

精选理由

这篇论文解决了表现强化学习中稳定性理论的关键问题,提出了新的收敛保证和假设条件。

该论文研究了表现强化学习中混合策略的稳定性问题。作者提出了局部混合稳定性和全局混合稳定性两个概念,并证明了在任意可能不连续的环境映射下,加权每状态Hedge动态能以O(1/√T)的速率将局部稳定性差距降至零。研究还展示了局部稳定性与全局稳定性之间的本质差异,并针对全局稳定性引入了比Lipschitz敏感性假设更弱的有限转移范围假设。

原文 · arXiv cs.LG

Local and Global Stability in Performative Reinforcement Learning

In performative reinforcement learning the deployed policy shapes the environment that generates the learner's future data, and the natural solution concept is a performatively stable policy that is optimal in the environment it induces. Existing convergence guarantees rely on Lipschitz sensitivity assumptions on the environment map $π\mapsto (P_π, r_π)$, which are hard to verify and fail in settings such as multi-agent best-response dynamics. We instead study stability for mixtures of policies, and show that the resulting picture is fundamentally different from performative prediction, where randomization removes the need for any sensitivity assumption. We distinguish local mixed stability, an occupancy-weighted first-order relaxation that we show is equivalent to stationarity, from global mixed stability, which certifies against arbitrary deviating policies. Our first result is that a weighted per-state Hedge dynamic drives the local stability gap to zero at an $O(1/\sqrt{T})$ rate for an arbitrary, possibly discontinuous, environment map, both with exact and with trajectory feedback. The two notions genuinely differ: we exhibit an instance where local stability is achieved exactly but every mixture has global stability gap bounded away from zero. For global stability we introduce a bounded transition range assumption, strictly weaker than Lipschitz sensitivity, under which unweighted per-state Hedge converges up to a floor of $O(γε_P/(1-γ)^3)$, and we prove a matching-in-$ε_P$ lower bound of $Ω(γε_P/(1-γ))$ under trajectory feedback, so this floor is unavoidable. Finally, we extend both notions to $n$-player performative Markov games, obtaining local stability with no assumption on the joint environment map or game structure, and global stability for performative Markov potential games.