arXiv 论文:代码文档对 coding agent 解决真实 issue 并无帮助
Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
一群研究者认真测了文档到底能不能帮 coding agent,结果在 10 个仓库上翻车,还公开了打分基准和提示词,做 agent 的人该看看这个负面结果。
一篇 arXiv 论文提出 roundtrip 基准,用「根据文档重新生成的代码能否通过原始测试」来为代码描述打分,发现描述的完整性而非长度决定保真度。作者以该基准为优化信号,找到一条能达到完全保真度的描述生成提示词,并成功泛化到未见文件。随后在 2 个模型家族和 10 个代码仓库上测试「更好的文档能否帮 agent 解决真实 issue」,结果显示否定:当源码在场时,静态精简文档和检索上下文都不如 issue 本身。论文同时公开了基准、优化器和文档起作用的边界条件。
Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
We investigate whether natural-language documentation helps coding agents resolve software issues, and we build the tools to construct and evaluate it. We introduce a roundtrip benchmark that scores code descriptions by whether code regenerated from them passes the original tests, and show that completeness, not length, drives a description's fidelity. Using the benchmark as an optimization signal, we discover a description-writing prompt that reaches full fidelity and generalizes to unseen files. We then test the hypothesis that motivated the work: that better documentation helps an agent resolve real repository issues. Across two model families and ten repositories, and against a positive control confirming that our evaluation can detect a genuine improvement, we find that it does not. When the source is present, neither static compact documentation nor retrieved context beats the issue alone. We report this negative result together with the benchmark and the optimizer, and we characterize the boundary at which documentation helps.