arXiv 论文用 Stackelberg 博弈分析 AI 安全防御的充分条件
Defensive Sufficiency in a Stackelberg Model of AI Security
一篇把 AI 安全防御写成博弈论的论文,算出了什么时候修 bug、做红队才划算,还证明修复快不代表被攻陷概率低。
论文研究自动化测试、人类红队测试和事件响应构成的安全反馈机制何时足够保护 AI 系统。作者证明,当每次未解决的攻击有持续被发现概率、修复有效且更新不破坏既有保护时,有限输入构成的攻击面被防御的概率为 1。论文还推导完成时间上界,并构建防御者主导的 Stackelberg 博弈模型,刻画威慑攻击的最低成本投入及均衡情形。数值实验显示更快修复可缩短被攻陷持续时间,但不一定降低被攻陷概率。
Defensive Sufficiency in a Stackelberg Model of AI Security
Feedback from automated testing, human red teaming, and incident response can strengthen an AI system's defenses when discovered failures lead to effective repairs. We study when this feedback process provides sufficient protection and when investing in it is economically worthwhile. We begin by showing that an attack surface composed of finite number of inputs is defended with probability 1 if every unresolved attack has a persistent chance of discovery, repairs are effective, and subsequent updates preserve earlier protection. We derive completion-time bounds and extend the analysis to growing attack surfaces, repairs that generalize across related attacks, and multiple discovery mechanisms. These results distinguish eventual protection against each fixed attack from complete protection at a single time. We then formulate a defender-led Stackelberg game in which the defender invests in proactive discovery and reactive repair, anticipating the attacker's choice of search effort. We characterize the least-cost allocation that deters attack and the equilibrium regimes in which the defender funds neither capability, one capability, or both. Numerical experiments illustrate these regimes and show how faster repair can reduce compromise duration without reducing compromise probability.unified theory of performance limits in generative language models.