论文精选

Harness感知蒸馏框架提升小模型智能体性能

Really nice paper on harness-aware distillation for small language model agents. (bookmark it) It'...

精选理由

这篇论文提出了一种新颖的蒸馏方法,让小模型智能体更好地利用harness信息,性能超越大模型教师。

研究人员提出Harness-Aware Distillation框架,用于训练小语言模型智能体。该方法通过对比有和无harness信息时的教师模型输出,训练学生模型优先选择使用harness信息的行动。在ALFWorld基准测试中,学生模型在未见任务上达到63.4%成功率,超越8B参数教师模型和47.0%的最佳基线。该方法还使模型能够逃脱59.7%的停滞情况,显著高于基线的46.8%。

原文 · elvis

Really nice paper on harness-aware distillation for small language model agents. (bookmark it) It'...

Really nice paper on harness-aware distillation for small language model agents. (bookmark it) It's a super interesting distillation framework for small language model agents that are deployed with a harness. In simple terms, you run the big model with and without harness info, then train the small model on cases where its action changes. Researchers show that adding the harness to on-policy distillation raises how often the student uses harness information on ALFWorld (65.7% to 73.1%) but leaves success flat (43.1% to 43.5%). Their method, Harness-Aware Distillation, queries the same teacher with and without the harness information and trains the student to prefer the action chosen with it. A filter drops pairs whose preferred action contradicts the harness records. The method uses no task rewards or success labels. The student reaches 63.4% on unseen ALFWorld tasks, compared with 47.0% for the best baseline, and exceeds its 8B teacher. It also escapes 59.7% of stalls, while the baselines stay near the untrained student's 46.8%. Paper: arxiv.org/abs/2610.02858 Chat with Paper: academy.dair.ai/papers/harness… 💬 12 🔄 4 ❤️ 21 👀 2562 📊 17 ⚡