论文用六维文化回应标准审计 SimpleQA 与 Chatbot Arena
Whose Facts Count? A Culturally Responsive Audit of LLM Evaluation Benchmarks
SimpleQA 全部题目只认英文档案,Arena 七成多提问是英文,这篇论文把评测偏差算成了具体数字。
一篇 arXiv 论文借用文化回应性评估(CRE)框架,用六个维度审计了 OpenAI 的 SimpleQA(4,326 题)和 LMSYS Chatbot Arena(600 段对话)。SimpleQA 的全部 4,326 道题都以英文档案文献作为验证依据,且仅一位标注者对哥伦比亚建国年份的偏好就占全部题目的 2.70%,虚增了全球南方覆盖的表象。Chatbot Arena 中英文提示词占 76.3%,而国际电信联盟(ITU)估计英文用户仅占全球网民的 25.9%。论文构建的 50 题反向基准测得 SimpleQA 的文化回应性缺陷均值高出近三倍(Cohen's d = 1.01),并据此提出一套可操作的 CRE 评估框架。
Whose Facts Count? A Culturally Responsive Audit of LLM Evaluation Benchmarks
LLM benchmarks function as evaluation instruments, informing decisions that affect education, labor, and public services worldwide. Drawing on Hood, Kirkhart, and Hopson's culturally responsive evaluation (CRE) frameworks, this paper applies a six-dimension CR rubric to audit OpenAI's SimpleQA (N = 4,326 items) and the LMSYS Chatbot Arena (N = 600 conversations). Every SimpleQA question requires English-language archival verification as its evidentiary basis. A single rater's preoccupation with Colombian founding dates accounts for 2.70% of items, inflating the appearance of Global South coverage. English-language prompts constitute 76.3% of Arena conversations, against an International Telecommunication Union (ITU)-estimated 25.9% share of global internet users. A 50-item counter-benchmark scored a mean CR deficit nearly three times lower than SimpleQA (Cohen's d = 1.01). The paper proposes a practical CR evaluation framework. These are structural validity failures, not incidental measurement problems, with direct consequences for communities whose knowledge traditions these instruments were not built to see.