论文

SCA:强化学习中GUI代理的空间信用分配

SCA: Spatial Credit Assignment for Reinforcement Learning of GUI Agents

精选理由

清华团队提出SCA方法,解决GUI代理强化学习中空间信用分配问题,在专业领域和动作预测基准上表现优异。

SCA方法通过利用采样点击的屏幕坐标来改进组相对信用分配。在所有采样点击都失败的情况下,SCA会根据点击到标注目标的距离进行排序。研究显示,SCA在专业领域GUI定位上有所改进,并在大多数动作预测指标上取得了强化微调模型中最强的结果。

原文 · arXiv cs.AI

SCA: Spatial Credit Assignment for Reinforcement Learning of GUI Agents

GUI agents automate tasks on digital devices by grounding language instructions in visual interfaces. Existing group-relative reinforcement learning improves GUI action prediction by comparing the rewards of multiple responses sampled from the same GUI state. However, binary evaluation treats spatially different failed clicks as identical and provides no relative signal when all sampled clicks fail. To address these limitations, we propose Spatial Credit Assignment (SCA), which uses the screen coordinates of sampled clicks to refine group-relative credit. Specifically, SCA predicts each held-out response's reward from the other responses in groups containing both successes and failures, then uses the prediction residual to adjust credit. When all sampled clicks fail, SCA instead orders them by distance to the annotated target. These spatial references are used only to construct the training update; the deployed policy remains unchanged. We evaluate whether this correction improves the policy update itself by comparing its error and directional alignment with the exact return gradient in a controlled synthetic study. Across GUI grounding and offline action-prediction benchmarks, SCA improves grounding across professional domains and achieves the strongest results among reinforcement-fine-tuned models on most action-prediction metrics, with consistent gains across the reported GUI suites.