线性循环记忆足以提炼机器人空气曲棍球策略
Linear Recurrent Memory Suffices to Distil a World-Model Policy for Robot Air Hockey
MIT研究证明,在机器人控制任务中,简单的线性记忆机制配合非线性表示学习就足够了,不一定需要复杂的非线性循环动态。
研究人员在模拟空气曲棍球防御任务中测试了DreamerV3教师模型。当跟踪丢失时,具有记忆能力的模型表现优于无记忆策略。使用64维状态的纯线性循环模型(k=0)在五种匹配种子下与GRU基线和教师模型性能相当。线性模型比GRU需要更少的循环参数和计算量。
Linear Recurrent Memory Suffices to Distil a World-Model Policy for Robot Air Hockey
Does memory-dependent control need nonlinear recurrent dynamics? We study simulated air-hockey defence under temporary loss of puck tracking. A DreamerV3 teacher outperforms a memoryless policy under tracking loss, while resetting the teacher's recurrent state sharply reduces performance, which demonstrates that the task requires memory. We distil this teacher into compact recurrent policies with a 64 dimensional state, with a combination of a diagonal linear recurrence and an optional rank-$k$ nonlinear innovation while retaining nonlinear observation encoders and action heads. Across five matched seeds, the purely linear recurrent model ($k=0$) matches both the GRU baseline and the teacher throughout the tested range of tracking loss. Increasing nonlinear innovation rank providing no measured benefits. This result is obtained on a fresh test split, which will be only opened after all models and analyses are frozen. The linear model requires fewer recurrent parameters and less computation than GRU, but performs comparably. These results suggest that, for this memory dependent control task, nonlinear representation learning around a simple linear memory mechanism can be sufficient, and that nonlinear recurrent dynamics are not necessarily required. These conclusions are limited to the simulated task, teacher, state dimension, and blackout horizon considered here, and to policies whose observation encoder and action head remain nonlinear.