A new training method called FloWright teaches every agent in a multi-agent AI workflow to improve together, not just the one agent that designs the workflow.
Multi-agent AI systems split complex jobs, like processing documents, building slides, reading charts, writing code, solving math, and handling finance tasks, across several specialized agents. Earlier training methods only fine-tuned the agent that builds the workflow plan, leaving the agents that execute each step untouched, even though a bad result could come from any of them. FloWright tackles that gap with a hierarchical reward system that figures out which role to credit or blame from a single pass-or-fail outcome, letting one agent self-improve or several co-evolve without extra models, labels, or runs. The researchers also built a companion method, DataWright, that hardens existing datasets by turning single-agent tasks into tougher, multi-step workflow challenges.
This matters because most agentic AI training today looks like a team where only the manager gets coaching while the staff stay frozen. Co-evolving multiple agent roles produced bigger gains (5.03 percent) than training one agent alone (2.83 percent), with the biggest jump hitting 7.41 percent on small open models, suggesting coordination between agents is a bigger lever than simply swapping in a larger model.
The results so far come from small open models on benchmark tasks, not live products, so whether this approach holds up inside the messier, higher-stakes workflows that AI vendors are busy selling to enterprises remains an open question.