AI/ ai · reinforcement-learning · llm-agents · machine-learning-research

New Training Method Grades AI Agents Step by Step

T2SPO estimates an AI agent's distance to success at each step, not just at the finish, and beats the standard method in tests on two benchmark tasks.

Researchers have found a way to tell AI agents how close they are to finishing a task, not just whether they finished it.

The method, called Trajectory-to-Step Policy Optimization (T2SPO), trains language-model agents using reinforcement learning. Normally these agents only get a reward at the very end of a task - success or failure - which gives little hint about which earlier moves actually helped. T2SPO instead looks back at the agent's own past successful attempts and uses a small statistical model to estimate, at every point in a new attempt, how many steps likely remain before success. Comparing that estimate from one step to the next gives the agent extra, step-level feedback during training, on top of the usual end-of-task signal. Researchers tested 1.5-billion and 7-billion parameter models on two benchmark environments - ALFWorld, a simulated household-tasks setting, and WebShop, an online-shopping one - and found T2SPO beat GRPO, a widely used reinforcement-learning baseline, on overall task success.

The sparse, end-only reward problem is one of the main things holding back agent training. Most existing fixes require someone to hand-write intermediate rewards for each specific task, which does not scale. T2SPO instead pulls that signal automatically out of the agent's own successful runs, with no task-specific reward engineering required.

Two benchmark environments and models no bigger than 7 billion parameters is a modest testbed. Whether this step-level credit assignment holds up on messier, real-world agent tasks - the kind with longer horizons and less forgiving failure states - is still unproven.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →