论文

LLM生成代码的库相关问题研究

Understanding and Mitigating Library-Related Issues in LLM-Generated Code

精选理由

这篇论文分析了LLM生成代码时常见的库错误问题,并提出了一种能显著减少这些错误的方法,对AI编程助手开发者很有价值。

研究人员分析了100个LLM生成的代码文件,发现84%包含至少一个库相关错误。这些错误包括不正确的导入路径、缺失导入、幻觉库、已弃用库使用和未使用导入。研究团队提出了一种代理方法,在300个真实世界Python框架任务上测试,包括LangChain和AutoGen,在GPT-5、DeepSeek-V3等五个模型上减少了38.1%-54.6%的库相关错误,提高了16%的代码正确率。

原文 · arXiv: DeepSeek

Understanding and Mitigating Library-Related Issues in LLM-Generated Code

Software practitioners increasingly rely on Large Language Models (LLMs) to generate code that integrates external libraries. However, LLMs often produce incorrect library usage, such as invalid imports, outdated API calls, and hallucinated dependencies, leading to compilation or runtime failures that reduce the reliability of AI-assisted software development. In this paper, we propose an agentic approach to mitigate libraryrelated errors in LLM-generated code. More specifically, we first conduct an exploratory study to characterize the library-related issues produced by LLMs. Our analysis of 100 LLM-generated code files reveals that 84% of generated files contain at least one library-related error, with recurring patterns including incorrect import paths, missing imports, hallucinated libraries, deprecated library usage, and unused imports. Based on these findings, we design an agentic approach that integrates task analysis, documentation grounding, code generation, and automated validation to improve library usage during code synthesis. We evaluate our approach on 300 code generation tasks derived from realworld implementations of rapidly evolving Python frameworks, including LangChain and AutoGen, across five LLMs: GPT-5, DeepSeek-V3, Qwen3, Mistral, and Llama 3. The results show that our approach consistently improves code generation quality across all evaluated models, reducing library-related errors by 38.1% - 54.6% and increasing code correctness by up to 16%.