论文

X-Reset:用人类手部演示做重置,解决灵巧操作强化学习探索难题

X-Reset: Scaling Object-Centric Reinforcement Learning via Cross-Embodiment Resets

精选理由

伯克利这边搞了个 X-Reset,把人类手部演示只用来当 RL 重置点,不用模仿,20 个物体三种机器人本体都能训出通用抓取策略,还能 sim-to-real 直接迁移。

arXiv 论文提出 X-Reset 框架,将人类手-物交互演示经运动学重定向为带噪声的机器人状态,过滤仿真中不稳定的状态后作为 RL 训练的重置分布。策略仅依赖物体状态和目标,配通用物体中心奖励,无需任务定制奖励或模仿人类动作。实验在 20 个物体、三种本体上训练出通用策略,包括 22 自由度机械手配两种机械臂和一个平行夹爪。该策略可扩展训练物体数量、泛化到未见物体,并能从仿真零样本迁移到真实机器人。

原文 · arXiv cs.LG

X-Reset: Scaling Object-Centric Reinforcement Learning via Cross-Embodiment Resets

Reinforcement learning (RL) in simulation can train dexterous manipulation policies without robot demonstrations, but training a single generalist policy with task-agnostic rewards faces a severe exploration problem: approaching, grasping, and reorienting diverse objects with many degrees of freedom is difficult to discover from scratch. Prior works make exploration tractable with high-quality robot demonstrations, per-task reward shaping, or by restricting policies to narrow modes of behavior. We propose X-Reset, a framework that instead resolves exploration with human hand-object demonstrations. Rather than imitating or tracking retargeted human motion, X-Reset kinematically retargets hand-object states to noisy robot states, filters out states that are unstable in simulation, and samples the remainder as resets during RL training with general-purpose object-centric rewards. The resulting policy depends only on object state and goal, with demonstrations entering training through the reset distribution. We show that X-Reset trains generalist policies on 20 objects across three embodiments---a 22-DoF hand on two different arms and a parallel-jaw gripper---and resolves the exploration challenges of RL from scratch. X-Reset scales with the number of training objects, generalizes to unseen objects, can learn from imperfect hand-pose estimates, and transfers behaviors zero-shot from sim-to-real.