Underwater C3-JEPA:面向 ROV 打捞的跨视角世界模型
Underwater C3-JEPA: An Object-Centric Cross-View World Model for ROV Salvage
一篇把世界模型用到水下机器人打捞的论文,不用接触传感器就能预测抓取时物体状态变化,真实水下视频也验证了,做机器人方向的可以看看。
论文提出 Underwater C3-JEPA,一个以物体为中心的多视角预测世界模型,用于近场重载水下 ROV 打捞任务。它无需接触传感器,仅凭多相机 RGB 观测和载体控制信号,在潜空间预测抓取交互中任务物体的状态演化。模型通过 held-out-view attention 融合跨相机证据,用 weak binding 低成本绑定目标与夹爪,并用 SIGReg 强化几何表征。实验显示其表征向下游探针迁移的任务相关信息显著多于无重建潜空间基线,且预测器保持轻量,可用于 MPC 候选评估和想象 rollout 的行为智能体训练。真实水下视频验证表明,同一架构能恢复被屏蔽相机视角下的物体状态并优于 persistence 预测。
Underwater C3-JEPA: An Object-Centric Cross-View World Model for ROV Salvage
We present Underwater C$^{3}$-JEPA (cross-view, control-conditioned, context-extended), an object-centric multi-view predictive world model for near-field heavy-load underwater ROV salvage. Without contact sensors, it predicts in latent space how the task-object state evolves through contact interaction and under the hydrodynamic lag of the vehicle, from synchronized multi-view RGB observations and vehicle control signals. C$^{3}$-JEPA encodes multi-camera observations into task-object and context tokens, fuses cross-camera evidence through held-out-view attention, and directly predicts future states conditioned on control. Weak binding anchors the target and gripper at low annotation cost, while SIGReg sharpens the geometric representation. Experiments show that the learned representation transfers substantially more task-relevant information to downstream probes than a reconstruction-free latent baseline, while keeping the predictor lightweight. The resulting predictive interface supports model-predictive-control (MPC) candidate evaluation and imagined-rollout behavior-agent training. Validation on real underwater video shows the same architecture recovering a withheld camera's object state and staying ahead of persistence, so the recipe transfers beyond simulation.