论文精选

SkillSonar:给智能体加运行时安全检查,攻击成功率从48.2%降到10.4%

精选理由

恶意skill会装的时候装乖,用的时候才下手,这篇论文的SkillSonar在GLM-5上把攻击成功率砍到10%左右,讲清了怎么防。

一篇论文指出,恶意 agent skill 可能在安装时表现正常,等智能体拿到文件、工具或凭证后才执行越权操作,因此仅靠安装前扫描不够。论文提出 SkillSonar,一个在运行时审查敏感操作的安全 skill,可决定放行、收窄、重新规划或先询问用户。在 GLM-5 上测试,已知攻击类型的攻击成功率从 48.2% 降到 10.4%,未见风险类型从 60.6% 降到 11.5%。实验还发现,只安装该 skill 效果有限,必须显式要求智能体行动前查阅它。论文建议三层防护:安装前扫描、执行时检查、保留权限和沙箱等硬性保护。

原文 · rohanpaul_ai

A malicious agent skill can look safe when installed and turn dangerous only during a real task, so this paper argues that agents need runtime safety checks, not just pre-install scanning.

The problem is timing: a bad skill can wait until the agent has access to useful files, tools, credentials, or external services before pushing it beyond what the user actually asked for.

The paper proposes SkillSonar, a safety skill that checks sensitive actions while the agent is working and decides whether to allow them, narrow them, replan, or ask the user first.

On GLM-5, it cut attack success from 48.2% to 10.4% on familiar attack types and from 60.6% to 11.5% on unseen risk families.

A crucial result: simply installing the safety skill was much weaker.

The agent had to be explicitly told to consult it before acting.

overall, the paper says scan skills before installation, check their actions during execution, and still keep hard protections like permissions and sandboxing underneath.

– arxiv. org/abs/2609.01487

Title: "Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents"