评估LLM智能体在知识冲突中的认知谦逊
Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict
这篇论文揭示了AI智能体在知识冲突中的行为模式,对构建更可靠的AI系统有重要参考价值。
研究人员提出通过认知谦逊(EH)评估智能体,包含识别、解决和升级三个维度。研究通过知识冲突场景评估了四种智能体,发现高任务准确性与高认知谦逊并不对应。模型级干预可提高EH,但常以牺牲任务准确率为代价。
Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict
When retrieved evidence contradicts an agent's prior beliefs, does it revise its answer, acknowledge uncertainty, or persist with an incorrect conclusion? Existing evaluations of agentic systems focus primarily on task success, offering limited insight into how agents handle such conflicts. We propose to evaluate agents on epistemic humility (EH): the agent's willingness to recognize, act on, and communicate uncertainty during task execution. We operationalize EH through three trajectory-level behavioral dimensions: Identify, Solve, and Escalate (ISE). Through knowledge conflict, situations where the backbone language model's parametric knowledge contradicts the evidence it encounters, or where two contextual sources disagree, we evaluate two conflict settings: (1) controlled conflict and (2) naturally occurring conflict during multi-step agentic execution, each paired with matched no-conflict controls. Evaluating four agents, we find that higher task accuracy does not necessarily correspond to greater epistemic humility: some high-accuracy configurations recognize conflicts during execution but do not communicate unresolved uncertainty in their incorrect final answers. Trajectory-level analysis further reveals that agents frequently detect conflicts in early steps of execution but fail to maintain or resolve them in later steps. Finally, we show that model-level interventions can improve EH, but often at the cost of task accuracy, suggesting that epistemic humility emerges from the interaction among the backbone model, agent harness, and evaluation environment.