AI/ ai-agents · llm-reasoning · machine-learning · arxiv

Researchers Teach AI Agents to Skip Unnecessary Reasoning Steps

A new training method called RACE cuts AI agent reasoning costs by teaching models to skip reasoning steps when earlier thinking still covers the action.

A new training method teaches AI agents to skip reasoning steps when earlier thinking already covers the next move, without needing a separate check to find out.

Researchers built RACE, short for Reasoning Adaptation through Cross-Turn Estimation, to solve a specific waste problem: most AI agents re-reason before every single action in a multi-step task, even when a previous reasoning step already justifies the next one. The method uses a procedure called LoGiC that tests, turn by turn, whether removing a reasoning step changes the likelihood of the action that follows it. If removing it barely changes anything, that step gets flagged as redundant. Those signals then feed into both supervised fine-tuning and reinforcement learning, so the agent learns directly when thinking is worth the tokens and when it should just act.

Chain-of-thought reasoning is the default way agents are built to be reliable, but it is also the most expensive part of running them at scale - more tokens, more latency, more cost per task. On four agent benchmarks, RACE-trained models cut that reasoning overhead substantially while matching or beating the performance of models that reason every turn, suggesting a lot of current agent reasoning is habitual rather than necessary.

It is a narrow fix to a real problem, and worth watching whether it holds up outside the four benchmarks tested before anyone builds a product on top of it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →