AI/ ai · multi-agent systems · token efficiency · research

Lightweight models could cut AI agent costs by up to 97 percent

A new framework lets cheap small models handle multi-agent coordination, cutting GPT-4o token use by up to 97.2% without hurting accuracy.

A new multi-agent AI framework offloads the boring coordination work to cheap models, and the savings are enormous.

Researchers built S1-MAS, a system that splits multi-agent AI collaboration into two tiers: a lightweight "System One" controller handles routine coordination, like picking tasks, assigning roles, and deciding when to stop, while a compact reader pulls relevant evidence from approved sources. The actual open-ended reasoning, the hard part, still goes to a capable large language model. That split lets the system adapt collaboration on the fly without any task-specific training. Tested across seven benchmarks against three existing multi-agent frameworks, AgentVerse, DyLAN, and SelfOrg, S1-MAS matched or beat their accuracy while cutting GPT-4o token consumption by 44.9% to 97.2% and trimming end-to-end latency by 37.8% to 93%.

Multi-agent LLM systems have a dirty secret: most of their token budget goes to bureaucracy, not thinking. Every time agents negotiate who does what or pass messages around, that's billed at frontier-model rates even though the decision itself is trivial. Separating cheap coordination from expensive reasoning is the kind of unglamorous architecture fix that actually makes agentic systems affordable to run at scale, which matters more for adoption than any benchmark score.

It is still a research paper, not a product, so the real test is whether this division of labor holds up outside curated benchmarks and against frameworks built after it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →