论文官方一手精选

苹果研究团队提出RLTL;DR自我改进方法

RLTL;DR: Self-Improvement by Internalizing Self-Generated Feedback

精选理由

苹果团队的新方法让AI在没有成功案例的情况下也能自我改进,特别适合解决超难任务。

苹果研究团队提出RLTL;DR方法,解决强化学习中自我改进难题。该方法让代理在每次失败尝试后,根据验证器输出生成TL;DR洞察作为反馈。该研究针对极难任务,无需教师模型或示例解决方案。RLTL;DR通过内部化自我生成的反馈实现持续改进。

图片来源 · Apple ML Research
原文 · Apple ML Research

RLTL;DR: Self-Improvement by Internalizing Self-Generated Feedback

The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RLTL;DR. After each failed attempt, we show the policy the verifier outputs and let it write its own feedback, in the form of a single TL;DR insight. The next rollout is conditioned…