模型多源确认

Anthropic:模型自我解释不可信

精选理由

Anthropic直接点出大模型自我解释的不可靠性,这对理解AI行为边界很重要。

Anthropic明确表示,模型对其自身推理过程的解释不能作为判断其行为原因的可信证据。这使得Anthropic难以准确评估各类故障的严重程度。该观点来自Anthropic官方声明。

原文 · rohanpaul_ai

Anthropic states plainly that the model’s own explanation of its reasoning can’t be trusted as evidence of why it acted, which is exactly why they can’t cleanly judge how severe each of these failures was. https://t.co/m0Fu2Ts2xE