UOPD:用不确定性感知干预改进多轮智能体的在线蒸馏
UOPD: Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn Agents
这篇论文给多轮智能体蒸馏提了个新思路:只在学生拿不准的步骤让老师出手纠正,WebShop 分数最多提升 15.8%。做 Agent 训练的可以看看干预时机的选择。
UOPD 是一种针对在线蒸馏(OPD)的不确定性感知干预方法,核心是用教师模型对学生动作的低置信度来筛选高不确定性步骤。在低不确定性回合执行学生动作并使用标准 OPD 损失,在高不确定性回合采样教师动作并让学生通过监督微调模仿。该方法采用自适应不确定性阈值来控制计划干预率。在 ALFWorld、WebShop、Search 等智能体任务上评估,UOPD 在 WebShop 分数上相比标准 OPD 最高提升 15.8%。
UOPD: Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn Agents
On-policy distillation (OPD) trains a student on its own rollouts using dense supervision from a teacher. In multi-turn environments, a mistake at a critical decision step can redirect the subsequent rollout toward poor outcomes. We use low teacher confidence on student actions to select high-uncertainty steps for correction. In a controlled ALFWorld study, a single teacher correction at a low-confidence step improves subsequent student behavior and task success, motivating selective intervention during distillation. We propose UOPD, an uncertainty-aware intervention method for on-policy distillation. At low-uncertainty turns, UOPD executes student actions and applies the standard OPD loss. At high-uncertainty turns, it samples and executes teacher actions and trains the student to imitate them through supervised fine-tuning, which minimizes forward Kullback-Leibler divergence in expectation. UOPD utilizes adaptive uncertainty thresholds to target a scheduled intervention rate. Empirically, we evaluate UOPD across a broad range of agentic tasks, including ALFWorld, WebShop, and Search, demonstrating its superior performance over OPD methods and their variants. UOPD improves WebShop score by up to $15.8\%$ relative to standard OPD.