A new study tests whether a stack of AI agents can beat a scripted adversary at negotiation, and the clearest win comes not from arguing better but from quitting and starting over.
The setup pairs a multi-agent system, built from a five-pillar written constitution, a four-tier structure of agent roles, and a recovery routine called Cognitive Annealing, against a deliberately rigid adversary called the Gatekeeper, whose acceptance rules are fixed regular expressions rather than a real decision-maker. Across repeated five-run trials, a bare baseline agent never found the right phrasing to satisfy the Gatekeeper, a version with the constitution alone unlocked it once, and the full agent structure unlocked it twice. Adding an LLM Monitor and a hard output gate did not raise that success rate, though the researchers note it left a clearer audit trail. When the system did unlock the Gatekeeper, it took one turn, 7 to 8 model calls, and about 15,000 tokens, 67 to 73 percent cheaper than baseline; failed attempts cost 17 to 38 percent more than baseline instead.
The sharper result comes from a separate trap test, where the Gatekeeper tries to deadlock the agents into complying with a request they're supposed to refuse. Letting the agents reason their way out with more LLM text failed in all five attempts, while a scripted move, wiping the agent's working memory and issuing one fixed refusal, escaped the trap in all five, at no extra model cost. Notably, four of the five LLM-written refusals had already been waved through by an LLM Monitor before a separate automated check caught all five as malformed.
It is a small, artificial testbed with a known answer, not a benchmark for how AI agents behave in the wild, but it is a pointed rebuttal to the idea that stacking more model calls and committees on top of an agent makes it safer: here, a dumb, deterministic reset beat every clever argument the agents could generate.