ProVer论文解决智能体RL信用分配问题
ProVer论文用LLM评判者解决智能体RL的信用分配问题,比GRPO方法提升近10%,论文已发布在arxiv。
ProVer论文提出使用LLM评判者选择检查轨迹位置,通过采样成功与失败轨迹的连续片段计算优势值。该方法在ALFWorld、WebShop和SearchQA基准上,相比GRPO方法,Qwen3.5-2B模型相对提升9.91%,Qwen3.5-4B模型提升7.12%。即使使用较小模型作为评判者,该方法仍然有效。
Good paper on credit assignment for agent RL. The main finding is that you want an LLM judge to choose where to check a trajectory, and the rollouts to decide how much credit that step gets. GRPO gives every token in a trajectory the same advantage, so the training signal cannot tell the decisive step from the rest. ProVer has a judge compare successful and failed rollouts and name the segment it thinks caused the difference. It then samples continuations from just before and just after that segment and uses the change in success rate as the segment's advantage. Across ALFWorld, WebShop and SearchQA, this gives relative improvements over GRPO of 9.91% for Qwen3.5-2B and 7.12% for Qwen3.5-4B. It still helps when the judge is a smaller model. Paper: arxiv.org/abs/2609.36178 Chat with Paper: academy.dair.ai/papers/targeti… 💬 1 🔄 0 ❤️ 0 👀 516 📊 1 ⚡