Researchers have found a way to tell AI agents how close they are to finishing a task, not just whether they finished it.
The method, called Trajectory-to-Step Policy Optimization (T2SPO), trains language-model agents using reinforcement learning. Normally these agents only get a reward at the very end of a task - success or failure - which gives little hint about which earlier moves actually helped. T2SPO instead looks back at the agent's own past successful attempts and uses a small statistical model to estimate, at every point in a new attempt, how many steps likely remain before success. Comparing that estimate from one step to the next gives the agent extra, step-level feedback during training, on top of the usual end-of-task signal. Researchers tested 1.5-billion and 7-billion parameter models on two benchmark environments - ALFWorld, a simulated household-tasks setting, and WebShop, an online-shopping one - and found T2SPO beat GRPO, a widely used reinforcement-learning baseline, on overall task success.
The sparse, end-only reward problem is one of the main things holding back agent training. Most existing fixes require someone to hand-write intermediate rewards for each specific task, which does not scale. T2SPO instead pulls that signal automatically out of the agent's own successful runs, with no task-specific reward engineering required.
Two benchmark environments and models no bigger than 7 billion parameters is a modest testbed. Whether this step-level credit assignment holds up on messier, real-world agent tasks - the kind with longer horizons and less forgiving failure states - is still unproven.