A research team has found a faster way to get an AI system to figure out where a voice is coming from in a noisy room.
The system in question is built as a group of AI agents, not one big model, with separate agents handling localization, source separation, and classification, then comparing notes with each other to fix mistakes in real time. Earlier versions corrected location errors by checking a single quality score at a time, and those scores bounced around so much from moment to moment that the search for the right spot was slow and unstable. The new method instead has the system evaluate quality across a whole batch of candidate locations at once, giving it a clearer map of where to look. The result is faster convergence, better accuracy, and more stable corrections of larger localization errors, with only a minor increase in how long the quality-checking agent itself takes to respond.
This is the same problem that sits underneath smart speakers, video-call noise suppression, and hearing aids: picking out one voice's location in a messy acoustic scene. Most efforts to fix that lean on bigger, more complex models. This result instead comes from restructuring what information the system feeds itself rather than adding compute, which is a useful data point against the industry's default instinct to just scale up.
It is still a lab result with real-time performance measured in a controlled setting, not a shipped product, so the trade-offs will look different once it meets an actual noisy kitchen or conference room.