AI/ ai · ai-agents · efficiency · research

A Multiple Choice Trick Speeds Up Small AI Agents

A new framework lets small multimodal AI agents pick from preset reasoning options instead of writing it out, cutting inference time by up to half.

A team of researchers just made small AI agents stop rambling before every move.

The new framework, called Selection-based Structured Reasoning (SSR), swaps free-form chain-of-thought for multiple choice. Instead of generating a fresh paragraph of reasoning at each step, the model picks from a fixed menu of pre-written reasoning candidates, scored by likelihood against the current context. That scoring runs in parallel using a shared cache, rather than one slow token-by-token generation pass. The team tested SSR on 2B and 4B parameter models across seven multimodal search benchmarks, using both reinforcement learning and supervised fine-tuning.

For small models, that rambling reasoning was mostly theater: it cost compute without improving the answer. SSR cut per-turn reasoning latency by over 90% and total inference latency by 28-54%, while matching the success rate of comparable search agents at the same scale. That is the kind of efficiency gain that matters for running agents on limited hardware, not a marginal benchmark bump.

It will not make a model think better - it just makes thinking cheaper, which for most real-world deployments is the harder problem anyway.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →