论文

ComputerSD:基于实时反馈的计算机使用智能体在线自蒸馏

ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents

精选理由

清华团队推出 ComputerSD,让计算机使用智能体能从实时 GUI 反馈中学习,比传统方法效果更好。

ComputerSD 是一种在线自蒸馏方法,通过将执行 GUI 转换的实时反馈转化为策略学习的指导。该方法在 OSWorld-Verified 基准测试中,通用 Qwen3-VL-8B-Thinking 模型表现优于仅基于结果的 GRPO 方法 1.9 个百分点,专业 EvoCUA-8B 模型则高出 4.1 个百分点。研究还验证了该方法在分布外设置中的泛化能力。

原文 · arXiv cs.AI

ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents

Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning signals through privileged rescoring, but directly applying it to CUA online training presents two challenges: fixed guidance may become misaligned with the student's current state, and guidance-induced probability shifts may conflict with step-level correctness. We introduce ComputerSD, an online self-distillation method for CUAs that converts real-time feedback from executed GUI transitions into guidance for policy learning. A fine-tuned GUI analyzer produces guidance and a step-level value score after each action; the guidance provides privileged context, while the score regulates the resulting OPSD signals. ComputerSD jointly optimizes token-level OPSD and trajectory-level GRPO in a fully asynchronous training framework. On OSWorld-Verified, ComputerSD outperforms outcome-only GRPO by 1.9 and 4.1 percentage points on the general-purpose Qwen3-VL-8B-Thinking and specialized EvoCUA-8B backbones, respectively. Evaluation in out-of-distribution settings further supports the generalizability of ComputerSD. These results demonstrate the effectiveness of learning from real-time feedback through online self-distillation for CUAs.