AI/ ai · multi-agent systems · llm research · arxiv

AI Agents That Learn From Peers Still Don't Beat Solo Learners

New research puts self-improving LLMs in populations and finds peer copying speeds up discovery but never beats solo learners per token.

Researchers found that letting AI agents watch and copy each other's strategies speeds up learning, but it still doesn't beat agents working alone.

The study tested populations of large language models that revise their own skill files and choose whether, when, and who to copy from, all sharing one token budget for search, observation, and action. Researchers first ran established social-learning algorithms on three LLMs and found the peer effect backfired: the agents earned less reward per token than solo learners, exploring too narrowly or running out of budget before they could act. They then let the models write and revise their own skills instead, and watching peers changed behavior, helping one model find useful skills sooner and another spend less on private search. Even so, neither group beat independent learners working at the same cost.

This matters because the AI industry is racing to stitch solo chatbots into multi-agent systems, on the assumption that more agents talking to each other automatically means more capability. The result here is a reality check: copying from peers made the population's learning more efficient, concentrating discoveries around whichever agent found something first, but it did not make the group more effective at solving the actual task. That gap between efficiency and effectiveness is the same one that's tripped up human organizations trying to scale by adding more meetings instead of more insight.

Multi-agent AI has sold itself on the promise that more heads are better than one; this study is a reminder that more heads sharing a token budget can just mean more mouths to feed.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →