Researchers say they've found a fix for AI agents that get confused when you change your mind mid-task: ask the model one big question instead of many small ones.
A new arXiv paper looks at tool-use agents, AI systems that split a task into sub-goals and call outside tools to complete each one. When a user revises a request partway through, the agent has to decide whether to keep, patch, or discard each already-computed sub-result. The researchers tested language models from two vendors and found that asking a model to make that keep/patch/discard call one sub-result at a time produced unreliable answers. But when the model was asked to classify the nature of the revision just once, with a separate deterministic layer applying that decision across all cached results, it matched what the paper calls the cost-optimal oracle on all three models tested.
The paper frames this as a lesson about agent architecture, not raw model capability: the same models failed at fine-grained, node-by-node decisions but succeeded at a single coarse one. The team also proves that no policy relying only on a sub-result's local context can be both safe and cost-optimal, meaning the ask-once, propagate-deterministically design isn't just a convenient trick. Across three test environments, it recovered the full possible savings, 43 percent cheaper than restarting from scratch, while getting every answer right.
It's a narrow, unglamorous problem, but it's the kind that decides whether AI agents are actually usable for real work. Anyone who has changed their mind mid-conversation with a chatbot and watched it either ignore the change or start over from zero will recognize exactly what's being fixed here.