论文多源确认精选

NVIDIA 研究:给多模态模型加工具会削弱拒绝有害请求的能力

精选理由

NVIDIA 这篇论文发现 Claude、Gemini、Qwen 这些模型一旦接上工具,挡有害请求的能力就掉最多近七成,做智能体安全评测的得看看。

NVIDIA 一篇被 NeurIPS 2026 接收的论文测试发现,多模态模型接入工具后拒绝有害请求的能力普遍下降,相对降幅最高达 68.7%,平均 17.7%。测试覆盖 Claude Opus 4.6 和 4.7、Gemini Agentic Vision、Qwen3.5-122B-A10B 等模型,基准包括 MM-SafetyBench、HoliSafe 和 VLSBench。作者归因于两点:工具输出填满上下文后掩盖了原始请求的有害意图,模型注意力也转向描述工具返回结果而非做安全判断。在最终回答前重新插入原始请求和图片可以恢复部分拒绝能力。只做纯聊天场景安全评估的团队可能会高估智能体的安全性。

原文 · elvis

Does giving a multimodal model tools make it worse at refusing harmful requests? New work from NVIDIA, accepted at NeurIPS 2026, says yes for every model it tested. Refusal failures rise by up to 68.7% relative, and by 17.7% on average. The drop appears in Claude Opus 4.6 and 4.7, Gemini Agentic Vision, Qwen3.5-122B-A10B, and agent-tuned open models across MM-SafetyBench, HoliSafe, and VLSBench. The authors trace it to two causes. Tool outputs fill the context and bury the original request's harmful intent. The model also shifts its attention to describing what the tools returned instead of making the safety decision. Re-inserting the original request and image right before the final response restores part of the lost refusals. If your safety evals run only in plain chat, they may overstate how safe your agent is. Paper: arxiv.org/abs/2610.03938 Chat with Paper: academy.dair.ai/papers/mllms-f… 💬 4 🔄 0 ❤️ 6 👀 1427 📊 4 ⚡

  • arXiv cs.AI10-07 17:51原文
  • marktechpost10-06 17:30原文
  • Google AI10-07 14:13原文
  • Satya Nadella10-07 18:17原文
  • rohanpaul_ai10-07 19:21原文
  • NVIDIA AI10-07 20:18原文
  • shao__meng02:41原文
  • The Business Times: Tech10-06 05:09原文
  • The Straits Times: Business10-06 06:01原文
  • kimmonismus10-06 14:18原文