论文

RV-ICL:让机器人智能体从演示视频中递归学习的免训练方法

Recursive Video In-Context Learning for Agentic Robot

精选理由

机器人智能体光记文字学不会“怎么做”操作,这篇论文把一条演示视频拆成层级结构按需查看,两个基准成功率都提了几个点,做具身智能的可以看看。

论文提出 RV-ICL,一种免训练方法,把演示视频组织成从任务关键帧到接触瞬间的层级结构,供 LLM 智能体按需读取。智能体规划时读取粗层信息,执行到抓取等子目标时只加载对应短片段,每任务只需一条演示。基于 RPent 的实验显示,LIBERO-PRO 成功率从 92.6% 提升到 96.5%,LIBERO-Plus 从 86.7% 提升到 95.8%。相比直接把整段视频塞进上下文,该方法解决了推理变慢和抓取接触细节丢失的问题。

原文 · arXiv cs.AI

Recursive Video In-Context Learning for Agentic Robot

LLM agents that orchestrate frozen vision-language-action (VLA) policies improve across episodes through text memory, which records what the agent did but not how the task is done. A demonstration video shows it, but fits poorly into an agent's context. The full video slows every turn, fixed keyframes lose the contact detail that decides whether a grasp holds, and what the agent needs shifts from the task's structure while planning to the frames around each contact. We introduce Recursive Video In-Context Learning (RV-ICL), a training-free method that turns a demonstration into a hierarchy the agent navigates rather than a prompt it receives. The hierarchy is built from the sub-events of the demonstration, such as grasps and releases. Its levels grow finer, from keyframes of the whole task to phases, moments and short clips, and are exposed through read-only tools. The agent reads the coarse levels before planning. During execution it re-enters the hierarchy whenever a step needs more detail and loads only the clip of its current sub-goal. One demonstration per task is enough. Built on RPent, RV-ICL raises success from 92.6% to 96.5% on LIBERO-PRO and from 86.7% to 95.8% on LIBERO-Plus.