论文

PAMORT:四足机器人用多目标强化学习运送无固定堆叠载荷

Transporting Unsecured Stacked Payloads with a Quadrupedal Robot via Multi-Objective Reinforcement Learning

精选理由

四足机器人背上没固定的纸箱走坡道台阶,成功率 0.850 比基线 0.675 高不少,做法是多目标 RL 加在线调权,做机器人的可以看看。

arXiv 论文提出 PAMORT 方法,针对四足机器人运送无固定堆叠纸箱时移动性能与载荷稳定性的权衡。该方法先训练以偏好向量调节运动与载荷稳定奖励的多目标基础策略,再冻结策略训练权重调节器,依据本体感知在线调整偏好。仿真中在包括未训练过的三箱堆叠等配置下,运输成功率持平或超过单目标基线。在 Unitree Go2 实机上零样本迁移到坡道和台阶,8 项任务平均成功率 0.850,基线为 0.675。

原文 · arXiv cs.LG

Transporting Unsecured Stacked Payloads with a Quadrupedal Robot via Multi-Objective Reinforcement Learning

Transporting unsecured payloads with legged robots over uneven terrain requires balancing locomotion performance and payload stability, since aggressive motion can destabilize the payload even when the robot remains stable. We study quadrupedal transportation of unsecured stacked boxes on an edgeless torso-mounted board without dedicated payload sensors or active carrier mechanisms. To address this trade-off, we propose Payload-Adaptive Multi-Objective Reinforcement learning for Transportation (PAMORT). PAMORT trains a multi-objective base policy conditioned on a preference vector that weights locomotion and payload-stability reward groups, then trains a weight adjuster on the frozen policy to adapt this preference online from proprioception. In simulation, PAMORT achieves comparable or better overall transportation success than a corresponding single-objective baseline across different payload configurations, including an unseen three-box stack, despite training only with two boxes. Real-world experiments on a Unitree Go2 demonstrate zero-shot transfer to slopes and steps at or beyond the training difficulty, with mean success rates of 0.850 for PAMORT and 0.675 for the baseline across eight tasks. These results demonstrate robust unsecured-payload transportation with online adaptation of the locomotion--payload trade-off from proprioceptive information.