Perplexity 用提示引导自蒸馏让模型从自身错误中学习
New research: We post-trained a Computer model to learn from its own errors using hint-guided self-d...
Perplexity 把自家模型改成会从错误里自我纠偏,A/B 测试里工具调用失败少了 21.2%,做法写得很细。
Perplexity 发布新研究,采用 hint-guided self-distillation(提示引导自蒸馏)方法对 Computer 模型做后训练,让模型从自身错误中学习。在一次线上 A/B 测试中,较晚训练的 checkpoint 相较较早版本将工具调用失败率相对降低了 21.2%。该方法面向减少智能体场景中的工具调用错误。
New research: We post-trained a Computer model to learn from its own errors using hint-guided self-d...
New research: We post-trained a Computer model to learn from its own errors using hint-guided self-distillation. In a live A/B test, a later trained checkpoint reduced tool-call failures by 21.2% relative to an earlier checkpoint. 💬 7 🔄 5 ❤️ 80 👀 12650 📊 18 ⚡