高UTD比率下价值函数崩溃可恢复
Rank Collapse Is Recoverable, Growing $|Q|$ Is Not: Out-of-Sample Early Warning for Value Divergence in High-UTD Soft Actor-Critic
这篇论文揭示了高UTD比率下价值函数崩溃的可恢复性,提出了基于早期增长的预警方法,能节省约10%的计算资源。
研究人员在Soft Actor-Critic算法中发现,提高更新-数据(UTD)比率会导致两种'塑性损失':表示崩溃和值函数|Q|增长。使用宽度为2048且无归一化的缩放批评器,崩溃是可恢复的。在第15k步时,被标记运行的|Q|值比标记水平低一百倍以上,但增长速率已能预测标记(Harrell's C和样本外AUC在Walker2d上为0.78,在Ant上为0.98)。LayerNorm批评器降低了这一速率。
Rank Collapse Is Recoverable, Growing $|Q|$ Is Not: Out-of-Sample Early Warning for Value Divergence in High-UTD Soft Actor-Critic
Raising the update-to-data (UTD) ratio breaks off-policy critics in two ways grouped as "plasticity loss": collapsing representations and growing value magnitude $|Q|$. We separate them in Soft Actor-Critic (SAC) with scaled critics (width 2048, no normalization). Collapse is survivable: at UTD ratio 16, HalfCheetah critics with most units dormant keep learning, and the training guard, which stops runs whose loss or $|Q|$ explodes, never flags them. Within one high-UTD SAC configuration, runs start close together, and how far a critic's $\log_{10}|Q|$ has climbed by step 15k, its early growth, ranks the runs by how soon the guard flags them. At 15k, a flagged run's $|Q|$ sits a median of over a hundredfold below its flag level, yet the climb's rate already orders the flags (Harrell's C and out-of-sample AUC 0.78 on Walker2d, 0.98 on Ant, at UTD ratio 4). Dormancy does not. Aborting on this rate saves about a tenth of held-out Walker2d compute and stays net-positive live. A LayerNorm critic lowers the rate, removes the flag on Walker2d at UTD ratio 4 and lowers the return.