DepthWorld:面向机器人操作的 3D 世界模型
DepthWorld: 3D World Model for Robot Manipulation
用深度监督修视频世界模型的 3D 一致性问题,还顺手把 DROID 数据集标定成 DROID-3D,做机器人操作的值得看。
DepthWorld 是一个基于 Stable Video Diffusion 的世界模型,通过空间潜在分块技术同时预测多视角 RGB 和深度,且不改动预训练的 VAE。团队设计了结合学习式双目深度与联合因子图的标定流程,从 DROID 数据集构建出 DROID-3D,提供稠密度量深度和重新标定的多视角外参,90% 的外部相机片段重投影误差低于 0.7 像素。在相同训练预算下,深度监督使 RGB 预测的 PSNR 比 RGB-only 基线提升 1.48 dB,同时输出可供下游几何推理的准确度量深度。
DepthWorld: 3D World Model for Robot Manipulation
World models offer a data-driven alternative to traditional simulators for robotics, with applications spanning policy evaluation, improvement, and planning. All of these uses depend on faithful 3D geometry, yet current video-based world models are trained on RGB alone and produce rollouts that look correct frame-by-frame but do not compose into a consistent 3D world. Closing this gap requires progress on two fronts: large-scale 3D supervision for manipulation, and an architecture that can absorb it without disturbing strong pretrained video priors. We introduce a calibration pipeline that combines learned stereo depth with a joint factor graph, pooling all episodes collected from the same physical robot to recover its shared kinematic parameters alongside per-scene extrinsics. Applied to the DROID dataset, this yields DROID-3D, a calibrated 3D dataset providing dense metric depth and recalibrated multi-view extrinsics (achieving <0.7 px reprojection error on 90% of episodes for external cameras). We then train DepthWorld, a Stable Video Diffusion-based world model that jointly predicts multi-view RGB and depth via spatial latent tiling, leaving the pretrained Variational Autoencoder (VAE) unchanged. Depth supervision improves RGB prediction itself by +1.48 dB PSNR over an identical RGB-only baseline at equal training budget, while simultaneously yielding accurate metric depth for downstream geometric reasoning.