AI agents that work step-by-step usually cannot learn from a mistake until the whole task is over, and a new study says that lag is costing them plenty.
A paper posted to arXiv (arXiv:2609.35911) describes StepLearn, a way to update an AI agent's knowledge while it is still mid-task rather than waiting for the episode to end. Instead of banking lessons only after a run finishes, StepLearn treats each informative step as a hypothesis, checks it against what actually happens next, and only turns it into a reusable rule once it holds up outside the episode where it was first noticed. Model weights never change; the agent instead builds up an external, verified set of rules it can draw on later. Across five rounds on the WebArena-Lite and ALFWorld benchmarks, the paper reports StepLearn success rates of 59.9% and 84.0% with GPT-5-mini, and 57.8% and 88.1% with Qwen3.5-35B-A3B, beating EvoTest, the paper's strongest baseline, by 2.2 to 12.7 percentage points.
The real news here isn't the score, it's the timing fix. Most test-time learning setups make an agent finish an entire episode before any lesson gets recorded, which works fine on a benchmark but is useless if you want an agent to correct course mid-task. Validating each guess before letting it guide future episodes is what keeps that speed from turning into agents that confidently repeat one-off flukes.
Whether that validation step holds up against messier, real-world tasks - rather than shopping-site and household simulators - is the open question benchmarks this tidy rarely answer.