A new system called LEMON is learning to manage teams of AI agents so they solve problems without burning excess compute.
Researchers built LEMON, short for Learning Executable Multi-Agent Orchestration via Counterfactual Reinforcement Learning, to design the blueprint for multi-agent AI systems: which agent handles which duty, how much processing capacity each gets, and how tasks depend on one another. Instead of judging the whole system's final output and vaguely reworking everything, LEMON edits individual pieces of that blueprint, a role, a capacity level, a dependency link, and measures how each edit alone changes the result. That gives the training process a way to credit or blame specific design choices instead of the system as a whole. Tested across six reasoning and coding benchmarks, including MMLU, GSM8K, AQuA, MultiArith, SVAMP, and HumanEval, LEMON posted the best average score among the multi-agent orchestration methods evaluated, though the paper reports that average rather than a win on every individual benchmark.
The more telling number is what the paper calls the accuracy-token trade-off: how much correct output you get per token an LLM agent spends generating it. LEMON improved that ratio, meaning it found orchestration setups that reached similar or better accuracy without proportionally more token spend. That matters because most multi-agent frameworks assume more agents and more back-and-forth automatically means better answers. This work is aimed squarely at the bill.
Nothing here is shipping in a product, and the paper doesn't publish per-benchmark head-to-head figures against named rivals, just the six-benchmark average. Worth watching, not worth rebuilding your agent stack over yet.