别只信排行榜:通用 LLM 评测的五大局限与任务专属评估
Do Not Trust the Benchmark: Limitations of General LLM Rankings and a Case for Task-Specific Evaluation
选模型前别光看榜单分数,这篇论文把通用排行的坑掰开讲了五种,还给了个众包评测的路子。
这篇 arXiv 观点论文列出通用 LLM 排行榜的五大局限:参评系统与公开版本不一致、外部评测存在商业利益依赖、基准饱和与数据污染、模型钻评分规则空子、总分与用户实际任务相关性有限。作者主张评测应披露被测配置、验证题目与任务完成的真实性、同时报告成本和执行耗时、并明确泛化范围。论文以众包评测平台 Isotanta 为例:扩大题池可提升任务覆盖,重复采样能提高估计稳定性,但两者都不保证有效性或个性化。核心论点是选模型要看它在目标任务上的表现证据,而不是通用榜单上的名次。
Do Not Trust the Benchmark: Limitations of General LLM Rankings and a Case for Task-Specific Evaluation
Benchmark scores increasingly influence the development, marketing, and selection of large language models (LLMs). Yet an overall score is interpretable only in relation to the system tested, the questions included, and the conditions of evaluation. This perspective examines five connected limitations of general LLM rankings: differences between evaluated and publicly available systems; commercial incentives and dependencies in external evaluation; benchmark saturation, defective tests, and data contamination; models exploiting scoring procedures; and the limited relevance of general scores to users' tasks. Documented cases illustrate why these problems require different responses. I argue for evaluation procedures that disclose the tested configuration, validate questions and successful task completion, report performance alongside cost and execution time, and make the scope of generalization explicit. I then discuss \textbf{Isotanta}, a crowdsourced benchmarking platform, as a practical example of contributed questions and repeated evaluation. A larger question pool may improve task coverage, while repeated sampling can improve the stability of estimates on that pool; neither guarantees validity or personalization. The paper distinguishes the platform's current shared ranking from proposed task-specific and user-provided evaluations. Its central argument is that model selection requires evidence about performance on the intended work, not simply a high position on a general leaderboard.