AI agents just got faster, thanks to a 0.6-billion-parameter model doing the busywork.
Researchers describe LEAP, short for Learning Efficient Action Proposals, a method for speeding up LLM agents without changing what they decide. Agents normally work one step at a time: reasoning through a task, picking an action, waiting for it to execute, then starting the next step - which is slow by design. Existing fixes borrow speculative-decoding tricks from text generation, using a second, usually large, model to guess the next move before the main model confirms it. LEAP instead trains a small 0.6B model specifically on the target model's own action sequences, so it agrees with the target's real decisions most of the time and makes agents up to 60 percent faster end to end, with no measured drop in task success.
The interesting part is not the speed number, it is the framework behind it. The researchers show that a round of action speculation only pays off when a small, cheaply trained drafter predicts the target well and the task has enough steps left to amortize the cost of drafting and verifying. That turns agent speedup from a trial-and-error hack into something engineering teams can estimate ahead of time, and the drafter can reportedly be trained online, without first collecting a stockpile of past agent traces.
Every agent accelerator pitch should now get the same question: is the drafter actually trained on this agent, or just borrowed and hoping for the best?