论文

用 NeMo Guardrails 约束 LLM 自动处置安全告警:注入召回率从 25.0% 提到 94.5%

Constrained-Action AI Remediation for SIEM/XDR via a NeMo-Guardrails Proxy

精选理由

安全团队看这个:论文用 NeMo Guardrails 给 SOC 的 LLM 加护栏,注入识别召回率从 25.0% 干到 94.5%,还放出了对抗测试语料,思路可以直接抄。

一篇 arXiv 论文提出面向 SOC(安全运营中心)的受限动作架构,让 LLM 在 SIEM/XDR 环境中做告警分诊与修复建议。方案分两层:SIEM/XDR 控制平面把 LLM 输出限制在封闭意图词表内,由端点代理执行模板化命令,并配参数校验器兜底;外层用 NeMo-Guardrails 代理包裹分析师 LLM,设置输入输出护栏。在自建的 SOC 对抗语料上,开箱即用的护栏把提示注入召回率从 25.0% 提升到 94.5%,误报率仅 0.1%。红队实测确认封闭词表和参数校验能在命令越过信任边界前拦截 LLM 失效模式;由于护栏延迟较高,作者建议人工在环或延迟执行,而非内联控制。

原文 · arXiv cs.AI

Constrained-Action AI Remediation for SIEM/XDR via a NeMo-Guardrails Proxy

Security Operations Centers (SOCs) for information technology and operational technology share one incident-response problem: a flood of correlated alerts and too few analysts. Large Language Models (LLMs) are increasingly proposed as reasoning engines that triage alerts and, in autonomous deployments, issue commands that block IPs, kill processes, or quarantine files on production hosts. This coupling introduces a new risk: a single adversarial alert can become a remote code path through the LLM's reasoning, leading it to recommend an action the SOC then executes. We present a constrained-action architecture with two coordinated layers: (i) a SIEM/XDR control plane that grounds remediation in correlated host events and confines the LLM's output to a closed intent vocabulary whose templated commands are executed by thin endpoint agents, backstopped by an argument validator; and (ii) a NeMo-Guardrails proxy that wraps the SOC-analyst LLM with input- and output-rail policies, evaluated out-of-the-box against a SOC-specific adversarial corpus we release. The stock proxy lifts injection recall from 25.0% to 94.5% at a 0.1% false-positive rate, and a live red-team exercise confirms that the closed intent vocabulary and argument validator contain the observed LLM failure modes before any command crosses the trust boundary. As an architectural fit (not yet a measured operational-technology deployment), the constrained-action property suits critical-infrastructure settings where a wrong remediation has physical, not merely operational, consequences. The loop is best run human-in-the-loop or delayed: the measured rail latency keeps inline control out of scope.