TAPDreamer 论文:对抗补丁可将世界动作模型成功率降到 0
TAPDreamer: Transferable Adversarial Patches for World Action Models
用一块贴纸大小的补丁就能让机器人世界模型成功率归零,还能跨模型迁移,做机器人安全研究的可以看看这篇攻击机制分析。
TAPDreamer 是一种针对世界动作模型的对抗补丁攻击方法,只需使用公开 encoder,无需查询目标策略输出。该方法用单个源任务的 6 帧数据,在约 6.5% 的输入面积上放置补丁,即可在 40 个 LIBERO 任务中将 FastWAM 成功率从 97.7% 降到 0.0%,在 50 个 RoboTwin 任务中从 90.8% 降到 0.0%。同一补丁还能跨模型迁移,使两种 DreamWAM 配置的成功率分别降到 2.1% 和 0.8%,Motus 降到 10.0%。研究结论指出,世界动作模型的防御必须保护共享视觉 encoder,只保护动作生成部分并不够。
TAPDreamer: Transferable Adversarial Patches for World Action Models
World models learn to predict how their environment will evolve, making them an important foundation for general-purpose robotic control. Yet world action models depend on camera inputs whose manipulation can corrupt the visual representations used across tasks and action policies. Existing attacks on these models optimize against the victim's actions or predicted futures and therefore require access to target-model outputs. In this paper, we propose an attack, TAPDreamer, against world action models that instead uses a public encoder alone to construct a fixed local perturbation that transfers across tasks and action architectures. TAPDreamer requires no target-policy queries. Our key insight is that interactions between patch-induced changes in attention weights and value vectors broadcast a nearly identical representation shift far beyond the patch footprint, and this shift remains stable across task observations. Guided by this insight, TAPDreamer uses six frames from one source task to maximize the global L1 distance between clean and patched encoder representations. In closed-loop evaluation, one frozen patch per benchmark, covering about 6.5% of the input, reduces FastWAM's success rate from 97.7% to 0.0% across 40 LIBERO tasks and from 90.8% to 0.0% across 50 RoboTwin tasks; matched random patches retain 81.5% and 79.2% success. The same patches reduce success to 2.1% and 0.8% on two DreamWAM configurations and to 10.0% on Motus. These results show that protecting downstream action generation alone is insufficient: defenses for world action models must also secure shared visual encoders against persistent local perturbations.