论文

VLM-Safe-RL:用 CLIP 信号改进安全强化学习,事故率从 31.6% 降至 19.4%

Vision--Language Signals in Constrained RL: Safety Gains Without Anticipation

精选理由

把 CLIP 接进 PPO-Lagrangian 做安全强化学习,MetaDrive Hard 上事故率降了 12 个点,但作者自己发现信号并不能预判碰撞,分析很实在。

论文提出 VLM-Safe-RL 框架,将冻结的 CLIP 信号通过奖励塑造和增强的乘子更新整合进 PPO-Lagrangian。在 MetaDrive Hard 基准上,灾难事故率从 31.6% 降到 19.4%。但 FormulaOne-L2 分析显示,CLIP 信号并不能提前预测碰撞,VLM 项对 Lagrange 乘子的影响也可忽略。结论是观测到的事故率下降是有条件的,并未发现碰撞预判能力。

原文 · arXiv cs.LG

Vision--Language Signals in Constrained RL: Safety Gains Without Anticipation

Safe reinforcement learning seeks policies that maximise task performance while satisfying safety constraints. In driving benchmarks, however, collision costs typically appear only at the time of collision, providing no advance warning of an approaching hazard. Frozen vision--language models can provide dense semantic feedback, yet it remains unclear whether their scores anticipate collisions and which component drives an observed safety improvement. Episodic cost can also favour policies that make little task progress. To address these gaps, we propose VLM-Safe-RL, a framework that integrates frozen CLIP signals into PPO-Lagrangian through reward shaping and an augmented multiplier update. On MetaDrive Hard, which combines the densest traffic with the largest map, the catastrophe rate falls from 31.6\% to 19.4\%. FormulaOne-L2 analysis finds no evidence that the CLIP signals anticipate collisions and shows that the VLM term has a negligible effect on the Lagrange multiplier. These findings show a conditional reduction in observed catastrophe rate without evidence of collision anticipation.