技能空间射击提升机器人自主策略
Skill-Space Shooting for Autonomous Robot Policy Improvement
MIT新方法让机器人通过可重用技能自主学习改进策略,无需人工演示每个纠正动作。
研究人员提出技能空间射击方法,利用基础模型指导探索可重用技能来改进机器人策略。该方法通过将成功试验转化为策略改进,实现了跨任务的自主策略提升。实验表明,该方法能持续改进自主行动策略,同时共享技能可减少新任务的教学需求。该方法使可重用技能成为纠正式监督的来源,实现了可扩展和可泛化的策略改进。
Skill-Space Shooting for Autonomous Robot Policy Improvement
Robots deployed in the physical world must be able to improve beyond their initial training as they encounter new situations and failures. For this improvement to scale across tasks, it must make effective use of experience without requiring human demonstration of each correction. Recent agentic systems offer a way to reduce this reliance on human effort by using foundation models to autonomously compose learned behaviors to complete tasks. Yet completing tasks this way does not itself teach a task policy to overcome its own failures; that requires turning these behaviors into learnable corrections for the policy. Our insight is that many such corrections are familiar short behaviors, or skills: they recur across tasks and describe actions that foundation models can reason about from a scene. We introduce skill-space shooting, which uses foundation model guidance to explore corrections through these reusable skills and turn successful trials into policy improvement. Real-world experiments show repeated improvement in policies acting autonomously, while skills can also be shared to reduce the teaching needed to improve on new tasks. By making reusable skills a source of corrective supervision, skill-space shooting enables scalable and generalizable policy improvement within and across tasks. Additional results and videos at https://skill-space-shooting.github.io.