研究显示推理模型的思考过程会消解部分偏见但制造更多新偏见
Does Thinking Help Fairness? Reasoning Tokens Resolve Some Biases but Create More
这篇 arXiv 论文拿三个 32B 模型做了思考前后对比,发现思考反而制造约 5 倍的新偏见翻转,做公平性评估的人该看看。
一项针对 QwQ-32B、DeepSeek-R1-Distill-Qwen-32B 和 Qwen3-32B 的消融研究检验了思考过程对反事实公平性的影响。在 Adult、COMPAS、Credit 三个高风险决策数据集上的全部 9 个(模型,数据集)组合中,思考过程新产生的反事实翻转数量是消解数量的约 5 倍,且新翻转出现在接近饱和的模型置信度下。论文提出 CDPG 指标追踪思考深度上的偏见演化,发现偏见随思考加深而传播和放大。Bias Transition Matrix 分析表明,这种不对称双重效应源于反事实配对预测从非思考到思考状态的联合转移。
Does Thinking Help Fairness? Reasoning Tokens Resolve Some Biases but Create More
Thinking in reasoning language models (RLMs) has been subject to debate on whether it resolves or amplifies bias. Prior works have shown competing conclusions in both directions. Using a within-model thinking-vs.-non-thinking ablation across QwQ-32B, DeepSeek-R1-Distill-Qwen-32B, and Qwen3-32B on three high-stakes decision tasks (Adult, COMPAS, Credit), we show that thinking has an asymmetric dual effect on counterfactual fairness: it both resolves counterfactual flips produced by the non-thinking baseline and creates new flips at near-saturating model confidence. In all nine (model, dataset) combinations, the created flips outnumber the resolved flips by roughly 5 times. To explain the effect, we treat the thinking trace itself as a measurable site of fairness change and study it through two dynamic instruments: 1) We propose Counterfactual Depth Probability Gap (CDPG) to track bias evolution along thinking depth, and observe that bias propagates and amplifies with thinking. 2) We also formulate the Bias Transition Matrix (BTM) to show how predictions of counterfactual pairs change from non-thinking to thinking, and find that the asymmetric dual effect originates in the pair-state joint transition.