论文

TAIP 引擎为 CI 流水线中的 LLM 安全审计提供持续保障

Continuous Assurance of Agentic Security Auditors for Software Delivery Decision Gates

精选理由

CI 里用 LLM 审代码不放心?这篇给出把审计证据版本化、策略一改就 1.1ms 内重算保障态势的工程方案,数据还挺扎实。

论文提出 Policy-Evidence-Execution 分离模式,通过 TAIP Assurance Engine 实现持续控制态势保障(CCPA)。评估基于未修改的 RepoAudit,在固定的 Python 空指针解引用基准上保留了 80 次执行记录,覆盖 gpt-4o-mini 和 gpt-4.1 两种模型配置。策略或证据变更后,TAIP 重新计算保障态势,观测到的最大策略到态势延迟为 1.1 ms。在 1,000 个独立保障上下文下,单工作线程的完整重算耗时最高 1.62 s,低于预设的 5 s Decision Gateway 预算。

原文 · arXiv: OpenAI

Continuous Assurance of Agentic Security Auditors for Software Delivery Decision Gates

Large language model (LLM)-based repository auditors are increasingly deployed as security controls within continuous integration (CI) pipelines, where their findings admit, block, or delay software changes. As Agentic Software Development Life Cycle (SDLC) Security Controls, their non-deterministic behaviour changes the evidence, while organisational risk appetite and jurisdictional or data-sovereignty policy change its interpretation. Point-in-time audits therefore cannot maintain current assurance for merge decisions. We propose the Policy-Evidence-Execution Separation Pattern, implemented by the Trustworthy AI Posture (TAIP) Assurance Engine and operated as Continuous Control Posture Assurance (CCPA). By separating policy from stable execution and binding admitted evidence to a versioned Posture Tree, the same assurance logic operates across models, environments, and policy profiles. We evaluate the approach using unmodified RepoAudit on a fixed Python Null Pointer Dereference benchmark. The retained evidence repository contains 80 RepoAudit executions across two OpenAI model configurations, gpt-4o-mini and gpt-4.1. TAIP recomputes assurance posture after policy, evidence, and model-context changes and is evaluated across increasing numbers of independent Decision Gateway contexts. The maximum observed policy-to-posture latency was 1.1 ms across three policy-class cycles in one execution. At 1,000 independent assurance contexts, full policy-triggered recomputation with one worker recorded a maximum aggregate refresh of 1.62 s, below the predeclared 5 s Decision Gateway budget. These single-host measurements concern assurance over retained evidence and exclude RepoAudit execution and provider inference.