论文

用陈述偏好经济学方法评估大模型:无需标准答案的有效性检验

Validity Without Ground Truth: What Stated-Preference Economics Offers the Evaluation of Language Models

精选理由

没有标准答案的问题怎么评测模型?这篇论文借经济学问卷检验的思路给了一套可操作方法,还实测了 6 个模型,思路挺新鲜。

arXiv 论文 2610.10506v1 提出把陈述偏好经济学中的有效性框架(内容、建构、准则有效性及激励相容等概念)用于 LLM 评估,解决没有标准答案的问题。作者用 Vossler et al. 2023 年发表的水质经济价值调查对 6 个模型进行测试,检验如需求曲线应向下倾斜、支付意愿应随商品范围和收入变化等经济学理论预测。结果分化明显:两个较旧模型在家庭收入 75,000 美元水平上未通过最基础的检验,两个最新模型通过了全部理论有效性检验,但在聚合有效性上出现分歧。论文强调,通过有效性检验只说明模型回答连贯,不代表正确。

原文 · arXiv cs.AI

Validity Without Ground Truth: What Stated-Preference Economics Offers the Evaluation of Language Models

Many of the questions now put to large language models have no correct answer to score against: what a policy is worth, which option a user should choose, how to weigh competing values. Stated-preference economics has faced this problem for decades. It judges survey responses without knowing the true value, through a framework of validity and related concepts: content, construct, and criterion validity, reliability, incentive compatibility, and consequentiality. We argue that this framework is a general method for evaluating language models, and we set out what each concept means for LLM evaluation. We demonstrate the approach using a published water-quality stated preference economic valuation survey (Vossler et al. 2023) administered to six models. In this economic application, the validity tests take the form of predictions from economic theory: demand should slope down, and willingness to pay should respond to the scope of the good and to income. The tests separate the models sharply. Two older models fail the most basic test at a household income level of \$75,000, and the two newest pass every test of theoretical validity we can score, but diverge on convergent validity. Passing validity tests shows that a model's answers are coherent, not that they are correct.