论文

LLM决策审计中评估工具影响比人口偏见更显著

The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits

精选理由

这篇论文揭示了LLM审计中评估工具设计比人口偏见更能影响结果,对AI公平性评估有重要启示。

研究人员测试了五个大型语言模型在招聘、贷款和医疗分诊场景下的表现。实验包含40,726个请求,仅申请人姓名不同。结果显示,36项预设对比中无一通过校正。模型在透明审计中几乎总是被识别,在相同内容比较中表现一致,对首位列出候选人的奖励效应超过任何测量到的人口偏见影响。

原文 · arXiv cs.AI

The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits

Whether a language model looks demographically biased can depend on how the audit asks its question. A charitable-aid benchmark reports that the same models favor minority applicants when rating requests one at a time and penalize some when ranking side by side. We test whether that reversal generalizes to hiring, lending, and medical triage: 40,726 requests to five models, applications differing only in the applicant's name, and a primary test fixed before collection. It does not. None of 36 planned contrasts survives correction. The rating advantage keeps its sign at roughly half the published size, and a precision extension bounds any hiring ranking penalty below the published effect, though the lending and triage ranking floors sit above that margin, so the exclusion is conclusive for hiring ranking and for rating in all three domains only. Planted disparities tracking their injected sizes and a directional replication on the original aid materials bound these nulls. The audit is livelier than the demographics: models recognize transparent audits nearly always, tie every identical-content comparison whether the varying detail is race or a hobby, and reward first-listed candidates as much as any demographic effect we measure. Audit verdicts reflect audit construction more than demographic bias.