Researchers built an AI team to answer health checkup questions, then checked whether it actually understood you better than one AI working alone.
The study tested 120 Korean-language queries that bundled two to four separate asks - like requesting a lab result, a lifestyle tip, and advice on which specialist to see in a single message - using synthetic checkup records. A single large language model handled requests end to end in one setup; in the other, a coordinator identified each intent, farmed it out to a specialized agent, and stitched the answers back together. Four separate LLM judges scored the multi-agent answers higher, with one key metric rising from 1.695 to 1.797. Two human reviewers preferred the multi-agent output in roughly two-thirds of head-to-head comparisons.
The catch is where the improvement actually came from. Gains were concentrated in usefulness, consistency, and remembering to answer every part of a compound question - not in medical accuracy or safety. Numerical accuracy improved under only one of four judges, medical safety showed no real difference, and critical failure rates were nearly identical: 15.0% for the single agent versus 13.3% for the team.
In other words, more agents made the answers better organized, not more correct. That is a meaningful distinction for anyone building a health assistant on top of an LLM: splitting the work helps with juggling, not judgment. It also cost more - 1.3 times the latency and twice the price - for a quality bump that one of four judges could not even confirm.