论文精选

MetaPermit:用元属性策略框架提升 AI 智能体工具调用的安全性

MetaPermit: Scalable and Auditable Access Control for AI Agents via LLM-Inferred Meta-Attributes

精选理由

智能体调工具乱授权是老大难,这篇用元属性加固定策略,比 CaMeL、IPIGuard 更稳还更一致,做 Agent 安全的可以看看。

MetaPermit 是一个面向智能体工具调用的访问控制框架,将语义推断与安全执行解耦:LLM 推断每次工具调用的元属性值,固定策略据此放行或拒绝,决策可审计。在 AgentDojo 和 AgentDyn 基准上覆盖 7 个任务套件和 5 种攻击方法,MetaPermit 的一致性比 LLM 驱动的授权高出 31%,任务完成度最高提升 109%。面对间接提示注入(IPI)攻击,它优于 CaMeL 和 IPIGuard,且未执行任何恶意工具调用。

原文 · arXiv: OpenAI

MetaPermit: Scalable and Auditable Access Control for AI Agents via LLM-Inferred Meta-Attributes

The rise of autonomous AI agents equipped with tools has introduced significant security risks, ranging from unintended tool misuse to adversarial manipulation through Indirect Prompt Injection (IPI) attacks. In practice, deployed agent systems such as OpenAI Codex and Claude Code protect tool invocations through a combination of coarse-grained permission rules and LLM-based judgments about individual proposed actions. Both components, however, have important limitations: static policies must anticipate possible user intents and therefore do not scale to open-ended tasks, while LLM-driven authorization supports dynamic decisions but produces inconsistent outcomes and remains vulnerable to targeted IPI attacks. To provide scalable and more consistent authorization, we propose MetaPermit, a policy-based tool access-control framework that decouples semantic inference from security enforcement. By analyzing agent-user interactions, we derive a compact, task-independent set of meta-attributes that capture the relationships among the user's intent, the execution context, and the proposed tool call. These meta-attributes allow MetaPermit to authorize tool use without enumerating user intents. At runtime, an LLM infers the meta-attribute values for each proposed tool call, while a fixed policy evaluates these values to allow or deny the call, making each decision auditable through the inferred values and the applied policy rule. We evaluate MetaPermit on the AgentDojo and AgentDyn benchmarks, across seven task suites and five attack methods, using two widely deployed open-weight LLMs. The results show that MetaPermit produces 31% more consistent authorization decisions than LLM-driven authorization and outperforms the state-of-the-art defenses CaMeL and IPIGuard in both task completion, with improvements of up to 109%, and robustness to IPI attacks, with no malicious tool calls executed.