Researchers have diagnosed a very specific AI agent bug: it never knows when to stop.
A new paper puts a name on it: LLM Parkinsonism, borrowed loosely from the folk notion that work expands to fill the time available. The authors argue the issue isn't just how language models predict the next token, but that a single self-conditioned loop handles everything: proposing actions, interpreting scope, judging progress, and deciding when to stop. Their fix, called Global Executive Control (GEC) v0.2, splits that loop apart, giving stopping authority to a separate, uncertainty-aware governance layer. In a 24,000-episode benchmark capped at 40,000 tokens per episode, a baseline agent hit 67.42% success; simply letting it choose among multiple candidate actions pushed that to 96.53%; adding GEC's governance held success at 96.57% while cutting average token use by 36.4%, from 19,782 to 12,574.
The real story here is cost, not correctness. Most of the accuracy gain came from candidate selection, not governance; GEC's contribution is efficiency, and in a world where agent runs are billed by the token, a 36% cut in wasted compute is the more sellable number than any success-rate bump. It's also a tacit admission that letting a model keep going until it feels confident is a bad default architecture for autonomous systems.
Worth noting: this is all mechanistic simulation, not a test on a live production model, so the eye-catching diagnosis is running well ahead of the cure.