论文提出加权质量指数,多维评估代码修复模型表现
Beyond the Leaderboard: Multi-Dimensional Evaluation of Dense and Mixture-of-Experts Models for Automated Program Repair
想给项目挑代码修复模型的话,这篇用 QI 指标把安全性、维护性和成本都算进去了,还发现 MoE 用更少激活参数打平稠密模型。
一篇 arXiv 论文提出 Weighted Quality Index(QI),基于 ISO/IEC 25010 模型,综合功能正确性、可维护性、安全性和生成效率来评估自动程序修复(APR)。实验覆盖 Qwen2.5-Coder 的 3B/7B/14B 三个稠密模型和 16B 参数的 DeepSeek-Coder-V2-Lite MoE 模型(2.4B 激活参数),在 40 个 QuixBugs 和 90 个 Defects4J 缺陷上以相同硬件本地运行。结果显示权重方案变化会改变模型排名,单一指标评估容易掩盖权衡。MoE 模型在正确性上与 7B/14B 稠密模型无统计显著差异(McNemar 精确检验),但激活参数少 3-6 倍。
Beyond the Leaderboard: Multi-Dimensional Evaluation of Dense and Mixture-of-Experts Models for Automated Program Repair
Automated Program Repair (APR) with language models is usually evaluated by whether a generated patch passes the test suite, which can hide differences in maintainability, security, and computational cost. We propose a Weighted Quality Index (QI), inspired by the ISO/IEC 25010 software quality model, that combines functional correctness, maintainability, security, and generation efficiency under configurable weighting schemes. We evaluate three dense Qwen2.5-Coder models (3B, 7B, 14B) and the 16B-parameter DeepSeek-Coder-V2-Lite Mixture-of-Experts (MoE) model (2.4B active parameters) on 40 QuixBugs and 90 Defects4J bugs, all run locally on identical hardware to control for infrastructure effects. Model rankings change with the weighting scheme, showing that single-metric evaluation can hide trade-offs. The MoE model shows almost no statistically significant difference in correctness from the 7B and 14B dense models (McNemar's exact test) while using 3-6 times fewer active parameters, whereas correctness increases significantly across the three dense scales. These results suggest that active parameter count can be a more informative lens than total parameter count for sparse code models.