Researchers built a way for AI agent teams to evolve their own workflows without redoing finished work.
The system, called Inherit-MAS, has a meta-model generate a workflow of specialized worker agents, each with defined roles and tool permissions, then a separate judge model scores how well an attempt performed and flags what went wrong. Instead of scrapping a failed run and starting over, the system keeps what worked, drops the unhelpful parts, and applies one targeted fix - that is "workflow inheritance." A second mechanism, "execution inheritance," skips rerunning any agent call whose exact inputs have not changed. Tested with GPT-4o-mini as the worker model, it reached 55.4% completion on the WorkBench benchmark and 49.7% F1 on HotpotQA's open-domain question answering test, beating three existing evolving multi-agent baselines - EvoAgent, EvoMAS, and TacoMAS. Swapping in Qwen3-32B workers produced the same ranking.
The benchmark scores are the headline, but the real story is cost. Skipping redundant reruns cut worker-level token usage by up to 34.6% and total token spend by up to 18.1% compared to rerunning the same system with that feature switched off. For anyone running agent pipelines at scale, that is the gap between a feasible product and a runaway bill.
It is still a research paper tested on narrow benchmarks, not something you can pip install today - and beating three baselines most readers have never heard of is a lower bar than it sounds.