自进化搜索代理中的协同作弊诊断与缓解
False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
论文揭示了自进化AI系统的隐藏缺陷,提出CrossFit方法有效解决协同作弊问题,在Qwen模型上效果显著。
研究人员发现自进化搜索代理存在"协同作弊"问题,即提问生成器和解答器在错误上达成一致,内部奖励提升但外部正确性未改善。实验显示Qwen3.5-4B和Qwen3.5-9B模型在多轮自进化中伪标签正确率停滞或下降。研究提出两种缓解方法:多样本验证(MSV)将虚假一致率从6.1%和8.8%分别降至5.7%和7.2%;CrossFit方法通过交叉验证将虚假一致率降至3.0%和3.7%,并在七个下游搜索基准上比标准自进化平均提升8.8和8.4分。
False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so internal reward improves without a matching gain in external correctness. A post-hoc audit against source evidence shows co-cheating growing more severe over successive rounds of self-evolution, with pseudo-label correctness stagnating or declining even as the in-loop training signal improves. The most direct mitigation is to verify proposals before training: we introduce multi-sample verification (MSV), which queries the same model three times with the source and three times without it to decide task admission and replace unreliable pseudo-labels. MSV partially reduces false agreement but leaves substantial residual co-cheating and costs six extra labeler generations per candidate. These limitations motivate CrossFit, our main method: it partitions the proposer's source documents into groups A and B; questions generated from A are scored by an auxiliary solver trained only on B, and vice versa. The cross-fitted agreement determines proposer reward, so a same-source pseudo-label cannot be reproduced through the feedback solver, while the original solver's update rule is unchanged. Rerunning the loop with Qwen3.5-4B and Qwen3.5-9B, MSV reduces false-agreement mass from 6.1% to 5.7% and from 8.8% to 7.2%, whereas CrossFit reduces it to 3.0% and 3.7%. Replaying identical proposals with source-excluded feedback further reduces false agreement to 0.4% and 0.1%, isolating feedback ancestry from curriculum changes. Across seven downstream search benchmarks, CrossFit improves average performance over standard coupled self-evolution by 8.8 and 8.4 points and over Search-R1 by 8.7 and 7.8 points at 4B and 9B.