论文72°

Perplexity 公开 Computer 智能体后训练方法,工具调用失败率降低约 21%

We're sharing new research on our post‑training approach, which teaches the Perplexity Computer agen...

精选理由

Perplexity 公开了他们训练 Computer 智能体的后训练细节,用 RFT 加自蒸馏纠正错误工具调用,实测失败率降了 21%,做 agent 的话值得看看他们怎么用真实用户会话。

Perplexity 分享了训练 Perplexity Computer 智能体的后训练研究。该方法让模型从真实用户会话中学习,模仿成功轨迹,并对整体成功但存在可避免错误的会话(如错误的工具调用)做显式纠正。技术路线上结合了拒绝采样微调(RFT)与提示引导的自蒸馏(hint-guided self-distillation)。在真实流量的 A/B 测试中,较新 checkpoint 相比早期 checkpoint 将工具调用失败率相对降低 21.2%。

原文 · Aravind Srinivas

We're sharing new research on our post‑training approach, which teaches the Perplexity Computer agen...

We're sharing new research on our post‑training approach, which teaches the Perplexity Computer agent to learn from real user sessions by imitating good trajectories and explicitly correcting avoidable mistakes like bad tool calls (even when the overall trajectory was successful). The method combines rejection sampling fine‑tuning (RFT) with hint‑guided self‑distillation, and it cuts tool call failures by about 21% in live A/B tests Perplexity @perplexity_ai New research: We post-trained a Computer model to learn from its own errors using hint-guided self-distillation. In a live A/B test, a later trained checkpoint reduced tool-call failures by 21.2% relative to an earlier checkpoint. 🔗 View Quoted Tweet 💬 10 🔄 5 ❤️ 105 👀 11454 📊 17 ⚡