VLA-Dreamer:用世界模型改进 VLA 机器人控制
Towards VLA-Dreamer: Refining VLA Behavior Using World Models
机器人方向的概念论文,想在 VLA 的视觉嵌入上训练世界模型,从而省掉海量模仿数据,还能做短期规划,思路挺有意思。
这是一篇概念论文,提出在 VLA(Vision-Language-Action)模型视觉编码器的嵌入空间上训练预测性世界模型。作者假设这些嵌入与动作相关,可用于预测未来,以检验 VLA 是否具备模拟真实世界动态的隐式世界模型。与标准世界模型在像素空间计算损失不同,该架构在嵌入空间计算损失,类似 JEPA。训练后的世界模型可通过给定目标图像采样 VLA 动作,用于短期规划任务,目标是降低 VLA 对大量模仿学习数据的依赖。
Towards VLA-Dreamer: Refining VLA Behavior Using World Models
Vision-Language-Action models (VLAs), while showing strong potential for robot control, require massive amounts of high-quality imitation learning data. Moreover, the absence of an explicit world model casts further doubt on their control capabilities. In this concept paper, we propose a novel architecture that addresses sample efficiency in VLAs by training a predictive world model on the embedding space of the VLA's vision encoder. We hypothesize that these embeddings are action-relevant and usable for future prediction. To this end, we propose using the suggested architecture to investigate how well these embeddings predict the future based on actions, as the inability to do so would mark a key limitation of VLA architectures: the lack of a non-lossy implicit world model to simulate real-world dynamics. The proposed architecture differs from the standard world model dynamics as the loss comes from the embedding space rather than the pixel space, similar to joint embedding predictive architectures. Furthermore, the trained world model can be utilized for short-term planning tasks by sampling VLA actions given goal images. We intend to examine the richness of vision embeddings in VLAs and reduce their high data requirements through a world model that can also generate plans during inference.