论文

TIDE:动态平衡教师蒸馏与强化学习的智能体训练新方法

TIDE: Teacher-Student Transition via Informative Distillation and Exploration for Agentic RL

精选理由

arXiv 上这篇 TIDE 讲怎么让小模型智能体训练时先跟老师学、再靠自己 RL 变强,比固定混合比例的做法更灵活,做智能体训练的可以看看。

TIDE 针对 GRPO 在小模型上早期探索受限的问题,改进了教师蒸馏(OPD)与强化学习的混合策略。该方法不使用固定混合比例,而是根据师生分歧趋势在训练过程中逐步将权重从蒸馏移交到 RL。在单轮交互内部,TIDE 依据相对动作价值和归一化分歧对蒸馏信号与 RL 优势进行逐轮调制。论文在多个基准和学生模型规模上的实验与消融验证了该自适应协调机制的有效性。

原文 · arXiv cs.AI

TIDE: Teacher-Student Transition via Informative Distillation and Exploration for Agentic RL

Effective multi-turn agents require interaction strategies that coordinate information gathering, actions, and feedback over long horizons. GRPO is a reinforcement learning algorithm used to train these agents, but sparse trajectory-level rewards limit early exploration in small models. Recent methods augment RL with on-policy distillation (OPD) from a stronger teacher. However, a fixed mixture assumes that teacher guidance and reward optimization should retain a constant relative role throughout training and across interaction turns. This assumption can fail at two scales. Globally, as training progresses, maintaining strong distillation pressure can constrain the model from moving beyond the teacher's capabilities. Locally, teacher--student disagreement identifies where the student departs from the teacher, but cannot tell whether that departure is exploration supported by better outcomes or low-quality policy drift. Our methodological insight is that teacher guidance and reward optimization should be dynamically rebalanced over training and jointly allocated across turns. We instantiate this insight in \tide. Globally, \tide uses the measured disagreement trend as a practical schedule signal, advancing an OPD-to-RL handoff when discrepancy reduction becomes slow but remains positive and progressively increasing the relative weight of RL. Locally, \tide jointly modulates teacher-guided and reward-driven updates: relative action value and disagreement prioritize the OPD signal, whereas relative action value supplies the RL advantage and normalized disagreement reweights it across turns. Coupled with the global handoff, \tide allocates stronger teacher guidance early and gives reward-driven updates greater relative weight later in training. Experiments across multiple benchmarks, student scales, and controlled ablations support the effectiveness of TIDE's adaptive OPD--RL coordination.