Long-WAM:扩展世界-动作模型上下文,实时机器人控制成功率提升至 78.7%
Long-WAM: Scaling the Context of World-Action Models
这篇论文把机器人模型的历史上下文拉到 19.2 秒,动态杯子堆叠做到 95% 成功,Pi0.5 一个都做不成,方法细节写得很清楚。
Long-WAM 是一个用于扩展因果世界-动作模型上下文的模型-系统框架,核心发现是视频基座采用自回归预训练时,更长历史才真正带来收益。在 RoboCasa GR-1 基准上,上下文从 0.0 秒增加到 19.2 秒,成功率从 63.3% 提升到 78.7%,而双向预训练初始化没有净增益。Long-WAM 在 LIBERO-Long、RoboTwin 2.0 和 DOMINO 上取得对比方法中的最佳结果,RTX 5090 上每个动作块耗时 107.4 毫秒。实际部署在 Unitree G1 和 YAM 上,动态杯子堆叠任务成功率达 95%,而 Pi0.5 和 Fast-WAM 在 20 次试验中全部失败。
Long-WAM: Scaling the Context of World-Action Models
Real-time robot control demands enough visual history to infer motion and task progress, but processing that history can delay action. We present Long-WAM, a model-system framework for scaling the context of causal world-action models under real-time control constraints. Our central finding is that access to history is not the same as using it: longer histories pay off far more when the video foundation is pretrained autoregressively (AR). We first learn causal prediction from robot and egocentric videos without action labels, then preserve this history-to-future structure during world-action adaptation. On RoboCasa GR-1, increasing context from 0.0 to 19.2 seconds raises success from 63.3% to 78.7%, whereas a bidirectionally pretrained initialization shows no net gain; robot-domain AR pretraining further raises peak success on GR-1 and LIBERO-Long. Long-WAM also achieves the best results among compared methods on LIBERO-Long, RoboTwin 2.0, and DOMINO. Streaming observation encoding, asynchronous execution, and hardware-specific acceleration enable deployment on RTX 5090, DGX Spark, and Jetson AGX Thor without dropping future prediction; on RTX 5090, each action chunk, including future-video latent prediction, takes 107.4 ms. Real-time deployment on Unitree G1 and YAM supports dynamic and long-horizon manipulation, including 95% success on dynamic cup stacking, where Pi0.5 and Fast-WAM succeed in none of 20 trials. As a memory-informed executor, Long-WAM also complements higher-level planning in composite tasks.