论文精选

EESD:用 Dirichlet 证据加权改进代码智能体自我纠错学习

How Much Evidence Should a Coding Agent's Self-Correction Carry? Adaptive Dirichlet Evidence for Self-Distillation

精选理由

一篇教代码智能体"怎么从自己的纠错里学到更多"的论文,在 DeepSeek/CodeARC 上把 Pass@1 从 15.0% 拉到 20.4%,做法挺有意思。

论文提出 Effective-Evidence Self-Distillation(EESD),把执行反馈的"相关性支持"与"证据质量"分开建模,用 Dirichlet 后验生成带不确定性惩罚的 KL 学习权重。在 DeepSeek/RunBugRun 匹配实验中,两种设置的 argmax 预测在全部 3,000 个样本上一致,但 EESD 的 NLL 更低。可观察样本从 1 增到 8 时,未来结果 NLL 降低 55.0-59.3%。在 DeepSeek/CodeARC 上做一轮纠错学习后,all-tests Pass@1 从 15.0% 升到 20.4%,配对 95% bootstrap 区间为 [+2.8, +8.0] 个百分点。

原文 · arXiv: DeepSeek

How Much Evidence Should a Coding Agent's Self-Correction Carry? Adaptive Dirichlet Evidence for Self-Distillation

Execution feedback lets coding agents revise programs and learn from their own corrections. A correction's learning weight should reflect both the transitions supported by its executions and the amount of evidence behind that support. We introduce Effective-Evidence Self-Distillation (EESD), which represents these quantities separately. Normalized execution relevance determines relative transition support and an effective pseudo-count mass; a Dirichlet posterior then produces an uncertainty-penalized weight for KL-anchored correction learning. Under a symmetric prior, changing mass preserves category ordering, and effective mass yields a supervised coefficient bounded by its matched fixed-mass counterpart. Across four model-domain history sweeps, increasing visible observations from one to eight reduces future-outcome NLL by 55.0-59.3%. At eight observations, effective mass achieves lower NLL than fixed mass in all four comparisons. In the primary matched DeepSeek/RunBugRun study, argmax predictions agree on all 3,000 examples, with the largest NLL gain under concentrated relevance. After one correction-learning round, DeepSeek/CodeARC all-tests Pass@1 increases from 15.0% to 20.4%, with a paired 95% source-bootstrap interval of [+2.8, +8.0] percentage points. The twelve-setting downstream evaluation establishes the model-domain scope of this update. These results show how separating evidence support from evidence mass changes probability estimation and correction learning in coding agents.