AI/ ai · robotics · machine-learning · research

AI Agents Learn How Many Steps to Take Before Replanning

A new policy called PACE teaches AI agents how many steps to take before replanning, boosting success while needing fewer decisions.

A new research method lets AI agents decide for themselves how long to act before stopping to rethink.

Researchers propose PACE (Policy-Native Adaptive Decision Timing), a vision-language-model-based policy that predicts both the next action and how many steps to execute before replanning, all within a fixed decision budget. They tested it on four long-horizon tasks - Sliding Puzzle, Sokoban, ALFWorld, and ScienceWorld - spanning fully and partially observable settings and both visual and text-based inputs. Across all four, PACE raised success rates by 3.1 to 15.6 percentage points while using fewer decisions on average than the baselines it was compared against.

Most AI agents today either lock in a fixed number of steps before checking back in, or use hand-written rules to decide when to replan, and neither approach adapts to what the agent actually runs into. Making that timing something the model learns, instead of a knob an engineer tunes in advance, could make long-running agents cheaper to operate and less prone to compounding errors in anything that requires many sequential actions, from warehouse robots to multi-step software agents.

It's a lab result on puzzle games and simulated kitchens, not a shipping product. But the problem it targets - knowing when to stop and think again - is the same one that slows down every agent asked to act independently for more than a few steps.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →