论文

因果表征学习中的块解耦:从可识别性理论到视觉状态估计

Block Disentanglement in CRL: Bridging Identifiability and Visual State Estimation

精选理由

一篇把因果表征学习理论落到机器人视觉上的论文,弱化干预假设还能免标注恢复物理状态,做具身智能的可以看看

这篇 arXiv 论文(编号 2610.06809)研究干预式因果表征学习(CRL)的块解耦问题。作者在明显更弱的干预假设下建立了可识别性保证,块结构由现实中可用的干预机制决定。随后该框架被用于机器人系统的视觉状态估计,直接从图像和视频恢复潜在物理变量,无需标注数据。实验显示即使假设被进一步违反,该方法在受控具身环境中仍然有效。

原文 · arXiv cs.LG

Block Disentanglement in CRL: Bridging Identifiability and Visual State Estimation

Causal representation learning (CRL) is the process of recovering causally-related latent variables from high-dimensional observations. As a label-free inference method, CRL is particularly attractive for applications where data labels are unavailable or impractical to obtain. While there has been significant progress in understanding the identifiability guarantees of CRL, such guarantees often hold under highly stylized assumptions, which temper the direct application to real-world problems. This paper has a two-fold objective for interventional CRL. First, it establishes identifiability guarantees for substantially weaker interventional assumptions, resulting in block disentanglement of the causal variables, where the block structure depends on the realistically available intervention mechanisms. Secondly, the block disentanglement framework is used for embodied visual state estimation, in which the objective is to recover the latent physical variables of a robotic system directly from visual data (images and videos) without labeled data. These two components are critically complementary. The block disentanglement theory delineates identifiability guarantees under weakened assumptions, and the application demonstrates that the resulting objective remains effective in a controlled embodied setting despite further assumption violations, providing a theory-to-practice bridge needed to translate the promise of label-free CRL into practical problems.