A new study finds that fancy multi-agent scaffolding adds little to nothing when autonomous coding agents tackle machine learning engineering tasks.
Researchers tested whether the elaborate harnesses used by today's top autonomous ML engineering (MLE) agents, things like multi-agent orchestrators and dedicated retrieval subagents, actually improve performance. They gave a minimal-harness coding agent, one with direct read, write and bash access to its environment, the same time budget and the same frontier LLM backbone as open-source state-of-the-art harnesses. Across a series of large-scale ablation studies, the stripped-down agent matched the complex systems on MLE benchmarks. The paper concludes the backbone model, not the surrounding machinery, is what drives results.
That is an awkward finding for a crowded field of teams selling orchestration layers, planning modules and subagent pipelines as the secret to better coding agents. If the scaffolding is mostly redundant, the real competitive edge sits with whoever has the strongest underlying model, not whoever built the cleverest wrapper around it.
It is a familiar lesson in AI tooling: the parts that look most impressive in a system diagram are often the parts doing the least.