PHRBench:评估大模型推理中纠正幻觉前提的行为基准
PHRBench: A Behavioral Evaluation of Post-Hallucination Reasoning in LLMs
arXiv 上新出的 PHRBench,测 18 个模型怎么在推理途中纠正幻觉前提,还发现提示特征能预测能否翻盘,做评测的可以看看。
论文提出 PHRBench,覆盖 4 个领域、18 个大语言模型,共 4820 个受控实例。基准通过幻觉顺从、幻觉规避和启发式纠正三个指标刻画推理轨迹,独立于最终答案正确性。结果显示成功纠正幻觉并得到正确答案的情况仍较罕见,且与推理路径上更频繁的信念更新相关。幻觉提示本身的特征对恢复成败有较强预测力,轻量预测器达到 0.847 的 AUROC。
PHRBench: A Behavioral Evaluation of Post-Hallucination Reasoning in LLMs
Hallucinated information can propagate through multi-stage LLM systems and become part of the context for subsequent reasoning. Existing studies of post-hallucination reasoning (PHR) mainly characterize changes in final outcomes and aggregate reasoning dynamics, leaving how models resolve hallucinated premises at the response level insufficiently understood. In this work, we introduce PHRBench, a controlled benchmark for behaviorally structured PHR across four domains and 18 large language models. PHRBench characterizes each reasoning trajectory independently of final-answer correctness through Hallucination Compliance, Hallucination Avoidance, and Heuristic Correction, and defines an insightful trajectory as successful correction that ultimately reaches the correct answer. Across 4820 controlled instances, we find that successful recovery remains relatively rare and is associated with more frequent belief updates along the reasoning trajectory. We further find that properties of the hallucinated prompt contain substantial predictive signal for successful recovery, with a lightweight predictor achieving an AUROC of 0.847. These findings provide a behavioral view of post-hallucination reasoning, characterizing how LLMs resolve erroneous context and when successful recovery is likely to occur.