论文

RLHF 用于人机协作的首篇综述:20 篇文献梳理与 VR 反馈时机实验

A Scoping Review and Experimental Study on Reinforcement Learning from Human Feedback for Human-Robot Collaboration

精选理由

机器人版 RLHF 怎么做?这篇综述梳理了 20 篇文献,还用 VR 实验证明让用户主动给反馈比系统定时收集更靠谱。

一项遵循 PRISMA 规范的范围综述筛选了 199 条记录,最终纳入 2020-2025 年间的 20 篇 RLHF 人机协作领域同行评审文献,声称是首个聚焦双向闭环设计的综述。研究归纳了多种反馈模态,指出人类反馈可嵌入 AI 训练的不同阶段,形成多步骤开发流程。为验证综述发现的关键缺口,作者开展了被试间 VR 实验,比较系统发起与用户发起两种反馈方式对机器人近身导航行为的影响。基于贝叶斯模型的分析显示,用户发起的反馈更能捕捉心理安全感(问卷)与物理安全(逆碰撞时间)指标,说明反馈时机直接影响反馈质量。

原文 · arXiv cs.AI

A Scoping Review and Experimental Study on Reinforcement Learning from Human Feedback for Human-Robot Collaboration

Human-Robot Collaboration (HRC) can facilitate mass customisation in Industry 4.0, with Reinforcement Learning from Human Feedback (RLHF) representing a promising approach for developing safe AI-based robots. Practical challenges remain regarding safety during AI development, human feedback quality, and bidirectional human-robot adaptation. We conducted a scoping review of RLHF in HRC systems, mapping methods that address these challenges. Following PRISMA guidelines, we screened 199 records and included 20 peer-reviewed publications (2020-2025) spanning multiple HRC domains. To our knowledge, this is the first review focused on the bidirectional, closed-loop design of RLHF. Our review found multiple feedback modalities enabling data collection in various feedback formats. Collected data can be integrated at different stages of AI training, resulting in a multi-step development process. Pilot experiments are commonly used to evaluate HRC systems based on both human and robot metrics. To empirically test a key gap identified in the review, we conducted a between-subjects VR experiment comparing system- and user-initiated feedback on robot proxemic behaviour for safe navigation. Using Bayesian models, we analysed the relation between the collected feedback and safety metrics: psychological safety (post-experiment questionnaire) and physical safety (inverse time-to-collision). Results show that user-initiated feedback captures perceived safety better than system-initiated feedback, indicating that feedback timing directly affects feedback quality. Our review and experiment findings show that RLHF relies on appropriate feedback methods to ensure AI safety in HRC, and future RLHF research should prioritise realistic HRC experiments evaluating the effects of feedback collection methods on relevant human and robot metrics.