Show-Harness:VLM智能体可操控机器人
Show-Harness: Just a VLM Agent Can Play Robots
Show-Harness让VLM能直接控制机器人,无需额外模型容量或特定预训练,零样本和微调两种方式都有效。
Show-Harness是一个具身接口,能让视觉语言模型(VLM)通过语义单元控制机器人。该系统可直接解锁闭源前沿VLM实现零样本机器人控制,也可通过少量GPU微调使小型开源VLM低成本部署。实验表明,配备Show-Harness的VLM智能体在任务、具身形态和环境方面表现出强大的泛化能力,超越了代表性智能体和VLA范式。
Show-Harness: Just a VLM Agent Can Play Robots
Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to "play" robots through a compact semantic interface linking intent to action. Show-Harness exposes discrete semantic action units that VLMs can naturally reason over, while embodiment-specific interpreters deterministically ground them into local robot actions, keeping the VLM directly responsible for fine-grained physical decisions. Through the same interface, Show-Harness demonstrates the feasibility of (1) directly unlocking closed-source frontier VLMs for zero-shot robot control, and (2) adapting small-scale open-source VLMs for low-cost deployment with just a few GPU-hours of fine-tuning. We further develop GUMI (GUI Manipulation Interface), which extends the same semantic action space to GUI-based demonstration collection, allowing humans and agents to "play" robots across embodiments without specialized teleoperation hardware. Extensive experiments show that Show-Harness-equipped VLM agents generalize robustly across tasks, embodiments, and environments, outperforming representative agentic and VLA paradigms. These results suggest that the right interface can unlock substantial embodied capability from foundation VLMs, without requiring additional model capacity or costly embodiment-specific pretraining.