Researchers just measured a specific way AI agents stumble: they ask for permission when the answer was already obvious, or barge ahead when they should have stopped to ask.
The team built a solver-grounded benchmark across four problem types - object allocation, meeting scheduling, apartment choice, and stable matching - to test whether models make the right call among three options: act, when every admissible preference agrees on one outcome; clarify, when multiple outcomes are feasible but none is shared; or propose a minimum-cost repair, when the request is simply infeasible. Matched pairs of scenarios kept the underlying problem identical while flipping whether intervention was actually needed, which let the researchers separate getting the label right from producing a usable action, question, or repair. The paper reports a recurring pattern across the benchmark rather than naming specific commercial models or publishing a leaderboard.
The interesting failure isn't that models miss uncertainty - it's that they sometimes find it correctly and still intervene unnecessarily, asking for clarification when an action was already justified. The paper also shows that how a response is required to be formatted changes not just the output's usability but which decision the model makes in the first place, which is an uncomfortable finding for anyone treating prompt formatting as a cosmetic detail.
If your AI scheduling assistant still fires off three clarifying emails before booking a conference room, there is now a benchmark that explains precisely why.