论文多源确认精选

Self-Evolving Defense:LLM智能体安全框架

Self-Evolving Defense: Continual Security Policy Learning for LLM Agents

精选理由

SED框架让LLM智能体无需重训练就能持续防御新型攻击,在DeepSeek、GLM和Kimi模型上都表现出色。

研究人员提出Self-Evolving Defense (SED)框架,无需重新训练模型即可持续学习安全策略。该框架在DeepSeek V4 Flash、GLM 5.2和Kimi K3三个开源模型上测试,在AGENTDOJO基准上将目标提示注入成功率降至0.42%,而最佳基线防御为3.7%。在HARMBENCH基准上,SED将X-TEAMING攻击成功率控制在7.8%,远低于最佳基线的35.2%,同时保持良性任务效用。

原文 · arXiv: DeepSeek

Self-Evolving Defense: Continual Security Policy Learning for LLM Agents

Large language models (LLMs) increasingly power agents that access sensitive information, use external tools, and modify software repositories. Although these capabilities offer substantial benefits, they also create security risks such as jailbreaks, prompt injection, and vulnerable code generation. Existing defenses often require retraining, fail to adapt to evolving attacks, or address only a single threat pattern. To address these limitations, we propose Self-Evolving Defense (SED), a training-free framework that distills harmful agent trajectories into reusable security policies without updating model weights. By retrieving relevant policies for future tasks, SED continually adapts to new attacks while retaining knowledge across attack scenarios. To evaluate the effectiveness of SED, we test it with three open-source models (DeepSeek V4 Flash, GLM 5.2, and Kimi K3) on eight benchmarks that span jailbreaks, prompt injection, and insecure code generation. SED lowers targeted prompt-injection success on AGENTDOJO to 0.42%, compared with 3.7% for the best baseline defense, and holds adaptive X-TEAMING attack success on HARMBENCH to 7.8%, more than four times lower than the best baseline at 35.2%, while preserving benign task utility.