研究显示 LLM 在工具调用场景下拒绝有害请求的阈值更高
Tool Mediation Alters Refusal Mechanisms in Large Language Models
做 Agent 的人该看看这篇:同一模型在工具调用时比纯聊天更容易放行有害请求,安全测试可能测不出这个坑。
论文《Tool Mediation Alters Refusal Mechanisms in Large Language Models》在多个开源权重模型上研究了工具调用对拒绝行为的影响。实验发现有害信息在模型内部表征中依然被强编码,且在对话与工具两种输入方式间可迁移。表征几何和神经元级分析显示两种模式下危害相关计算的分布不同:对话输入在较低感知危害水平即可触发拒绝,工具输入则要到更高的有效阈值才拒绝。逐步削弱拒绝计算时,工具模式下的拒绝在更低干预强度下被破坏。作者据此提出工具调用环境本身会降低模型对有害请求的鲁棒性,传统安全评测未必适用于 LLM 智能体。
Tool Mediation Alters Refusal Mechanisms in Large Language Models
Large language models (LLMs) are increasingly deployed with access to external tools, yet harmful tool-mediated interactions are less likely to be refused when compared to regular conversational ones. As this change in refusal behavior remains underexplored, we investigate its underlying mechanisms across a diverse set of open-weight language models. We find that information about the harmfulness of a request remains strongly encoded in the model's representations and transfers across conversational and tool-mediated inputs. Evidence from representation geometry and neuron-level analysis further indicates that the two interaction modes systematically distribute harm-related computation differently. Crucially, while conversational inputs can be refused at relatively low levels of perceived harmfulness, tool-mediated inputs remain permissive until harmfulness crosses a substantially higher effective refusal threshold. Moreover, tool-mediated refusal is also more brittle: progressively weakening the refusal computation disrupts tool-mediated refusal at lower intervention strengths than conversational refusal, even when benign capabilities remain intact. Together, our findings indicate that tool mediation does not simply reduce the internal perception of harm, but instead impacts its conversion into refusal. Overall, this suggests tool-mediated environments may intrinsically reduce robustness of models to harmful requests, and that conventional safety evaluations may not fully transfer to LLM agents.