A team of researchers just made small AI agents stop rambling before every move.
The new framework, called Selection-based Structured Reasoning (SSR), swaps free-form chain-of-thought for multiple choice. Instead of generating a fresh paragraph of reasoning at each step, the model picks from a fixed menu of pre-written reasoning candidates, scored by likelihood against the current context. That scoring runs in parallel using a shared cache, rather than one slow token-by-token generation pass. The team tested SSR on 2B and 4B parameter models across seven multimodal search benchmarks, using both reinforcement learning and supervised fine-tuning.
For small models, that rambling reasoning was mostly theater: it cost compute without improving the answer. SSR cut per-turn reasoning latency by over 90% and total inference latency by 28-54%, while matching the success rate of comparable search agents at the same scale. That is the kind of efficiency gain that matters for running agents on limited hardware, not a marginal benchmark bump.
It will not make a model think better - it just makes thinking cheaper, which for most real-world deployments is the harder problem anyway.