论文

DiSCo框架评估LLM文化偏好偏见

DiSCo: A Distribution-First Steering and Cultural Prior Evaluation Framework for Measuring Cultural Preference Bias in LLMs

精选理由

DiSCo框架首次系统测量了LLM的文化偏好偏见,发现提示引导无法解决文化偏见问题。

DiSCo是一个分布优先的评估框架,通过四级上下文梯度隔离大语言模型的默认文化先验。研究使用DiSCo-Bench数据集(304个项目,涵盖12种文化)评估了6个多样化的指令微调LLM。结果显示,默认先验高度集中,英国和美国文化占比约35%,仅代表12种文化中的2种。基于提示的引导持续扩大高资源文化与低资源文化之间的选择差距,而注入明确文化事实几乎无法改变分布。

原文 · arXiv cs.AI

DiSCo: A Distribution-First Steering and Cultural Prior Evaluation Framework for Measuring Cultural Preference Bias in LLMs

Large language models (LLMs) are increasingly deployed in globally used assistants, yet their default choices in culturally grounded everyday situations can systematically favour some cultures over others, affecting localisation, user trust, and equitable behaviour. Existing cultural benchmarks evaluate accuracy against a single "correct" answer, making it difficult to characterise an LLM's cultural preference prior when multiple culturally grounded responses are all valid; they also conflate default preferences with context-driven adaptation. We propose DiSCo, a distribution-first forced-choice evaluation framework that isolates default cultural priors and tests steerability via a four-level context gradient (C0--C3). Using DiSCo-Bench (304 items) derived from BLEnD spanning 12 cultures, we evaluate six diverse instruction-tuned LLMs. Default priors are heavily concentrated, with UK and US together absorbing approximately 35\% of all selections despite representing only 2 of 12 cultures. Most critically, prompt-based steering consistently widens the selection gap between high- and low-resource cultures, and injecting explicit cultural facts produces negligible distributional disruption, confirming that cultural preference bias cannot be resolved through prompt-based personalisation alone.