评估大模型助教对学生的教育公平性新基准发布
EduFair-Bench: Evaluating Pedagogical Fairness of LLM Tutors Across Student Demographics
想了解大模型助教是否真的公平?看这篇研究,他们用新基准测试了五个不同能力的模型,发现能力强的模型反而可能有更大的公平性差距。
研究人员推出了EduFair-Bench基准,用于审计大模型助教的教育公平性。该基准通过模拟不同背景(性别、移民背景、第一语言、社会经济地位)的学生与五个不同能力的LLM助教互动,评估助教的教学质量。研究发现,模型能力与公平性并不完全相关,语言和移民背景相关的偏见比性别和收入相关的偏见更大。
EduFair-Bench: Evaluating Pedagogical Fairness of LLM Tutors Across Student Demographics
Large language models (LLMs) are increasingly deployed as tutors, but it is unclear whether they support all students equally well. We introduce \textbf{EduFair-Bench}, a benchmark for auditing the pedagogical fairness of LLM tutors---whether tutoring quality varies systematically with student demographics. EduFair-Bench pairs a multi-domain question bank (mathematics, physics, chemistry) with a controlled simulation in which a fixed LLM student interacts with each tutor across nine demographic levels spanning four dimensions: gender, immigration background, first language, and socioeconomic status (SES). Tutoring quality is scored on five turn-level pedagogical metrics and four conversation-level dimensions, using an LLM judge validated against three-annotator consensus on 180 tutor turns. Bias is measured via paired Wilcoxon signed-rank tests and bootstrap effect-size confidence intervals. Two ablations (demographic cues conveyed through names; conflicting demographic information between tutor and student) disentangle tutor-driven from student-driven bias. Across five tutors, we find that model capability and demographic fairness are largely orthogonal: the smallest model is the most consistent while the four more capable tutors all exhibit wide demographic gaps with no clear capability-to-fairness ordering, pedagogy-specific RL training redistributes rather than removes bias, and language- and immigration-related cues produce larger gaps than gender- and SES-related cues.