论文

条件视觉定位问题诊断与改进

What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies

精选理由

MIT团队发现视觉干扰物如何导致机器人选错目标,并提出三种具体改进方法,在物理机器人上效果显著。

研究人员使用ACT模型系统研究了视觉干扰物对模仿策略的影响。实验通过控制颜色和形状相似性引入干扰物,发现失败主要集中在拾取和放置阶段。研究评估了三种干预方法:干扰物增强、阶段依赖注意正则化和基于外观的视觉提示,这些方法在模拟和UR3e物理机器人上显著提高了鲁棒性。

原文 · arXiv cs.AI

What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies

Visuomotor imitation policies can achieve high performance under in-distribution visual conditions yet fail when visually similar objects or receptacles are introduced. We study this behavior as a problem of conditional visual grounding: the visual target required for successful control changes with the manipulation phase and, in more complex tasks, with the observed task state. Using Action Chunking with Transformers (ACT), we systematically introduce distractor objects and receptacles with controlled color and shape similarity and localize failures to picking and placement. We find that distractor sensitivity is specific to both the type of visual similarity and the manipulation stage. Guided by this diagnosis, we evaluate distractor augmentation, phase-dependent attention regularization, and appearance-based visual prompting as complementary interventions for improving target selection while preserving spatial information required for control. These interventions substantially improve robustness in simulation and on a physical UR3e. We further examine the same failure pattern in a pretrained vision-language-action policy on a state-conditioned instrument-handling task, where the observed state of a medical instrument determines the correct destination. Together, the results show that visual distractors can cause incorrect object or destination selection even when the underlying manipulation skill remains intact, and that explicitly improving target selection can substantially recover performance across distinct visuomotor policy-learning regimes.