Security/ ai-security · game-theory · incident-response · red-teaming

A Game Theory Model for When AI Defenses Are Enough

A new paper models AI security as a defender-attacker game, showing when testing and incident response suffice to deter attacks and when they fail.

A new theoretical paper puts a number on something security teams usually argue about by gut feeling: how much testing and incident-response spending is actually enough to stop attacks on an AI system.

The paper models AI security as a Stackelberg game, where a defender sets spending on two activities, proactive discovery through testing and red-teaming, and reactive repair through incident response, before an attacker decides how hard to search for exploits. The authors show that a finite attack surface gets fully defended with probability 1 over time, as long as every unresolved vulnerability has some persistent chance of being found, repairs actually work, and later updates don't undo earlier fixes. They extend the model to attack surfaces that keep growing, repairs that generalize across related exploits, and multiple discovery methods. From there, the game-theoretic layer identifies the cheapest spending mix that deters an attacker outright, and separates the regimes where it pays to fund discovery only, repair only, both, or neither.

The sharpest finding for anyone setting a security budget: faster repair shortens how long a compromise lasts, but it does not by itself lower the odds that a system gets breached in the first place. That's a useful distinction, since discovery and repair get bundled together as "defense" even though only one of them is stopping attacks rather than cleaning up after them.

It's a clean theoretical framework, not a field study. The model assumes rational attackers and effective, lasting fixes, conditions that real-world AI deployments rarely guarantee outright.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →