荷兰书利润方法评估语言模型预测一致性,发现模型在概率预测上存在显著不一致,尤其在复杂逻辑关系下更明显。
该研究评估了语言模型在概率预测方面的一致性。研究人员通过构建线性程序计算荷兰书利润,衡量模型预测的不一致性。研究发现语言模型预测存在显著不一致性,事件间逻辑关系越丰富,不一致性越高。无关上下文细节可使不一致性提高一个数量级。
Dutch Books for Language Models
People increasingly use language models to support life decisions. Many such decisions involve a probabilistic forecast: How likely is a major life event, a natural disaster, or an economic outcome? Users of language models may implicitly trust that these forecasts fall out of a coherent world model. In this paper, we evaluate the coherence of language model probabilistic forecasts through a procedure that builds on a theorem due to de Finetti. We elicit forecasts from language models across events generated from stock returns data. We then use linear programs to compute the largest Dutch-book profit - the profit an arbitrageur could guarantee by betting against model-generated probabilities - which we use as a measure of incoherence. Our procedure does not require outcome labels, so we can evaluate coherence even in settings where outcomes are not observed or have not yet resolved. We find substantial evidence of incoherence in language model forecasts. Such incoherence increases when there are richer logical relationships between events, and irrelevant contextual details can increase incoherence by an order of magnitude. We conclude by discussing how alternative training strategies may improve probabilistic coherence.