AI/ multi-agent systems · ai research · benchmarks · arxiv

Researchers Test Letting AI Science Agents Coordinate Live

A new study lets AI research agents pick teammates and verify work mid-task, finding live selection helps but added bureaucracy backfires.

A new experiment asks whether AI research agents can swap jobs mid-task instead of following a fixed script.

Researchers built a system called Runtime Agent Coordination, or RAC, that lets AI agents working on scientific tasks pick teammates while a task is running, hand out scoped work assignments, and check each other's output using the artifacts they produce rather than sticking to a locked-in plan. The team tested RAC across three existing multi-agent "AI scientist" platforms - Agent Laboratory, EvoScientist, and ARK - using a benchmark called ResearchClawBench, while keeping each platform's own models, tools, and budgets intact. They ran a single-seed exploratory comparison of four setups: the normal fixed workflow, adding runtime communication, adding runtime agent selection, and stacking scoped contracts plus verification on top of selection.

Letting agents choose collaborators on the fly beat the fixed-workflow baseline on every platform tested, which is a real point against today's common practice of locking in the division of labor before a run even starts. But adding more coordination machinery - the contracts-and-verification layer - dragged average scores back down, with the effect varying by platform and in some cases landing below native execution.

Scoped contracts and built-in verification sound like the kind of bureaucracy that should make multi-agent systems more reliable, but here they mostly got in the way - a reminder that more process is not automatically more performance.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →