onPanda:用 token 级修正高效标注 LLM 与 Agent 的 on-policy 对齐数据
onPanda Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correc...
一篇教你怎么标注 on-policy 对齐数据的论文,思路是在 token 层面做修正,LLM 和 Agent 都适用,做对齐训练的可以看看。
onPanda 是一篇关于对齐数据标注方法的研究,核心是通过 token 级修正(Token-Level Correction)来标注 on-policy 数据。方法面向 LLM 和 Agent 两类场景,替代了传统对完整生成结果做标注的方式。论文已发布在 Hugging Face Papers 上,编号 2609.24。
onPanda Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correc...
onPanda Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction paper: huggingface.co/papers/2609.24… 💬 1 🔄 0 ❤️ 4 👀 972 📊 2 ⚡
- arXiv cs.LG09-21 17:56原文