用 SLM 做 Agent 每次工具调用的意图匹配验证
Toward SLM-based agentic task-tool intent matching
教你在 MCP 环境里用小模型逐条校验 Agent 的每次工具调用,用了 GRPO 加微调,安全审计方向可以看看思路。
一篇 arXiv 论文提出用 Small Language Models(SLM)作为任务-工具相关性分类器,对 AI Agent 的每一次工具调用独立验证是否贴合任务意图,弥补传统授权只能判断“允许调用”却无法判断调用是否合理的缺口。作者构建了一个多工具任务数据集,所需工具跨越不同 MCP 服务器。训练上结合了提示词优化、监督微调和基于 GRPO 的强化学习来专门化 SLM。该方案面向低延迟和本地部署场景,为 Agent 调用提供逐次执行层监督。
Toward SLM-based agentic task-tool intent matching
Tool-equipped AI agents use tool calls to access data and act on external systems. Horizontal growth of agentic systems increases the number of these interactions, and further motivates the need for automated, per-call oversight that can operate at low latency and/or on-prem. Conventional authorization schemes can determine whether an agent is allowed to invoke a tool, but cannot assess the agent's underlying cognition, specifically, whether the tool selection represents a logical, relevant step toward satisfying the intent of the task or not. Consequently, an allowed call may still deviate from the task's intent: a rogue agent might deviate the calls or nudge other agents to make a combination of calls that would not align with the intent of the task. Therefore, every call needs to be verified. In this study we investigate the applicability of Small Language Models (SLMs) to this purpose: an SLM functions as a task-tool relevance classifier that evaluates every selected tool independently against the assigned task and returns a relevance signal for downstream enforcement. Equipped with a novel dataset with multi-tool tasks whose required tools span distinct Model Context Protocol (MCP) servers, we used prompt-optimization, supervised fine-tuning, and reinforcement learning through GRPO to optimize and specialize SLMs.