A new benchmark tests what happens when AI models have to cooperate under pressure, and the results are not reassuring.
Researchers built GT-HarmBench, a set of 1,535 high-stakes scenarios structured around classic game-theory setups: the Prisoner's Dilemma, Stag Hunt, and Chicken. The scenarios were drawn from the MIT AI Risk Repository and cover situations like military escalation, election manipulation, and medical malpractice. Across 15 frontier models, the AI agents failed to choose the socially beneficial action in 38% of these high-stakes cases. The team also found outcomes shifted depending on how prompts framed the game and in what order options were presented.
Most AI safety evaluation still treats models as solo actors solving isolated problems, even though real deployments increasingly put multiple agents in the same room to negotiate, compete, or bargain. GT-HarmBench is a rare attempt to measure what happens when incentives conflict, not just whether a model refuses a bad prompt. The researchers also showed that simple game-theoretic interventions pushed socially beneficial outcomes up by as much as 18%, suggesting the problem is at least partly fixable with better framing rather than a fundamentally smarter model.
A 38% failure rate on decisions involving military escalation and election manipulation is not a rounding error. Alignment tested one model at a time, it turns out, does not automatically carry over once that model has to share the sandbox.