AI agents that stop to double-check their own moves perform better, but asking for that critique at every step is slow and expensive.
Researchers built SAG, short for Self-improving Agent with Gated critique, a framework that decides step by step whether an agent should pause for outside feedback. SAG reads two signals to estimate how confused the agent is about its next move: the overall spread of likely actions (entropy) and the gap between its top two choices (top-2 margin). Those signals approximate the Value of Information, essentially asking whether getting help right now is worth the cost. On the ALFWorld benchmark, task success jumped from 24.6% to 78.4% while token spending stayed close to a standard ReAct agent, a 3.1x gain in token efficiency, and a 7B-parameter agent paired with a lightweight 3B critic matched a 14B agent working alone.
That is a direct answer to the current default fix for unreliable AI agents, which is to scale up the base model or run critique on every single step and absorb the compute bill. SAG offers a cheaper third option: let the agent judge its own uncertainty and spend tokens only when it is actually stuck. For anyone running agents in production, where every extra model call shows up on an invoice, that tradeoff is the whole story.
This comes from one arXiv preprint tested on three interactive benchmarks, not a shipped product, so whether confidence-gated critique holds up on messier real-world tasks is still an open question.