论文

H-JEPA:分层世界模型端到端学习,提升长程视觉规划成功率

H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning

精选理由

LeCun 系 JEPA 路线的分层版本来了,AntMaze 成功率从 18% 干到 73%,做机器人规划和世界模型的值得看看这篇论文。

H-JEPA 是一种端到端训练动作条件 JEPA 层级结构的方法,每层在各自学到的潜空间中预测更远的时间尺度。规划自顶向下进行:顶层优化目标进度,每层的预测成为下一层规划器的子目标。在 Visual AntMaze 基准上,三层层级结构将成功率从 18% 提升到 73%,且规划计算量更少。结合逆动力学监督,该方法在 DROID 真实机器人视频上实现了更低计算成本下的离线规划保真度提升。

原文 · arXiv cs.LG

H-JEPA: End-to-End Learning of Hierarchical World Models for Visual Planning

Long-horizon planning with latent world models requires reasoning across timescales and levels of abstraction. Existing task-agnostic JEPA world models predict and plan at a single timescale or with multiple horizons in one shared latent space. We introduce H-JEPA, an end-to-end recipe for training a hierarchy of action-conditioned JEPAs in which each level predicts farther ahead in its own learned latent space. Planning proceeds top-down: the top level optimizes progress toward the goal, and each level's predictions become subgoals for the planner below it. When factors in the data evolve at separated timescales, higher levels discard fast, unpredictable detail and retain slower task-relevant state. Across four simulated navigation and manipulation environments, hierarchical planning improves over a flat JEPA; on Visual AntMaze, a three-level hierarchy raises success from 18% to 73% using less planner compute. Ablations attribute these gains to both temporal decomposition and higher-level goal representations. With inverse-dynamics supervision, the approach extends to diverse real-robot videos from DROID, where hierarchy improves offline planning fidelity at lower planner compute.