论文精选73°

微软研究:编码代理在代码理解上易出错

精选理由

微软发现编码代理在代码理解上容易出错,建议测试时重点评估阅读和比较代码的能力。

微软研究人员构建了CABRA,可生成合成编码任务并逐步提高难度。他们在6,840个任务上测试了8个LLM和6个代理,标记了每次工具调用类型。普通LLM随任务增长表现下降,但代理通过grep等工具保持近乎完美表现。在SWE-bench Verified基准上,阅读和分析调用次数比编辑行数更能预测代理失败,相关系数分别为-0.200和-0.159。

原文 · rohanpaul_ai

New Microsoft paper finds that coding agents trip up when they have to understand a lot of code, not when they have to edit a lot of it, so test them on reading and comparing code instead of diff size.

Microsoft researchers built CABRA, which generates synthetic coding tasks and raises 1 kind of difficulty at a time. They ran 8 LLMs and 6 agents on 6,840 tasks and labeled each tool call as reading, analyzing, searching, editing, or testing.

Plain LLMs got worse as tasks grew, but agents stayed near-perfect by using tools like grep. On SWE-bench Verified, the count of reading and analysis calls tracked agent failures better than lines edited, with correlations of -0.200 versus -0.159.