4D-HOF:前馈式流匹配实现4D手物交互重建
4D-HOF: Hand-Object Flow Matching for Feed-Forward 4D Interaction Reconstruction
一篇手物交互重建的新论文,4D-HOF 不从噪声生成,而是从基础模型的粗估计出发做流匹配,速度和稳定性都更好。
4D-HOF 是一个前馈框架,用于从视觉基础模型生成的粗略估计中重建 4D 手物交互。它训练条件流匹配模型,把基础模型给出的手物状态传输到交互流形上,以前馈方式修正平移、旋转和对齐误差。与生成式方法从随机噪声合成不同,它以基础模型的估计为起点,预测更稳定。重建过程中直接嵌入物理交互约束和 2D 观测证据作为测试时引导,无需事后单独优化。在域外基准上,4D-HOF 取得 SOTA 表现。
4D-HOF: Hand-Object Flow Matching for Feed-Forward 4D Interaction Reconstruction
Existing methods for 4D hand-object reconstruction often rely on costly per-sequence optimization, while generative approaches typically synthesize interactions from random noise, which can lead to unstable interaction prediction. We introduce 4D-HOF, a feed-forward framework that reconstructs 4D hand-object interactions from coarse but informative estimates produced by vision foundation models. Concretely, we learn a conditional flow matching model that transports foundation-model-derived hand-object states toward an interaction manifold, allowing the model to correct errors in translation, rotation, and alignment in a feed-forward manner. A key advantage of our generative formulation is that it naturally enables test-time guidance within the transport process. Rather than applying a separate post-hoc optimization after reconstruction, we directly steer the evolving generative states using physical interaction constraints and observed 2D evidence, allowing the reconstruction to be refined as part of the generative process itself. By training the generative model on diverse datasets, 4D-HOF generalizes robustly to challenging in-the-wild scenarios. Experiments on out-of-domain benchmarks show that 4D-HOF achieves state-of-the-art performance, producing more stable and accurate 4D hand-object reconstructions.