EmbodiedRSI:假设图驱动的自进化机器人学习框架
EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution
这个框架让机器人自己决定该做哪个物理实验,边试边改自己的代码和技能,真实机器人上也能零样本迁移,效果比基线高了快一倍。
EmbodiedRSI 是一个自进化 agentic 框架,用 Fast-Slow 双系统架构和 Hypothesis Graph 维护多个代码与技能假设,通过 Value-of-Information 实验选择来决定下一次物理探索。在 RoboCasa365 上总体成功率达 77.0%,Composite-Unseen 场景为 71.3%,而最佳基线只有 40.1%。在 LIBERO-Pro 上总体成功率达 86.8%。该框架还能零样本迁移到真实机器人,在多个挑战性任务上取得 71.3% 的总体成功率。
EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution
Robot foundation models provide strong visuomotor control, yet their performance can degrade when object positions or task instructions change. Further improvements often require post-training on substantial robot data, which can be costly to collect through methods such as teleoperation. Agentic harnesses can adapt around the model, but current self-evolving harnesses use robot trials inefficiently when deciding which code and skill changes to pursue. We introduce EmbodiedRSI, a self-evolving agentic harness that autonomously decides where to explore next and turns the resulting physical interaction into improved code and skills. EmbodiedRSI realizes this through a Fast-Slow Dual-System Architecture, in which competing code and skill hypotheses are maintained in a Hypothesis Graph. Value-of-Information Experiment Selection chooses physical experiments that can distinguish these hypotheses. Their outcomes guide Code-Skill Co-Evolution. The Slow System builds Hierarchical Memory, and Reward-Grounded Memory Learning selects effective memory according to their value for later Fast-System improvement. On RoboCasa365, EmbodiedRSI reaches 77.0% overall success and 71.3% on Composite-Unseen, compared with 40.1% for the best baseline. EmbodiedRSI also reaches 86.8% overall success on LIBERO-Pro. Beyond benchmark performance, EmbodiedRSI transfers zero-shot to real-world robot, achieving 71.3% overall success across multiple challenging tasks.