AI/ ai · generative-ai · audio · research

New AI Model Remixes Video Audio From Text Prompts

A generative AI model rebalances dialogue, removes unwanted sound, and reduces reverb in video using text and video guidance, researchers say.

An AI model can now clean up and remix a video's audio just by following text instructions.

Researchers describe the system, called Spot, Separate, and Enhance (SSE), in a paper posted to arXiv on September 25, 2026 (arXiv:2609.29169, project page at sse-ai.notion.site). SSE takes a video's soundtrack and rebalances it, strips out audio sources you don't want, and dials down reverb - guided by a mix of video frames and written descriptions of what should change. The team also built a new dataset, DegradedMix, layered on top of an existing audio remixing benchmark called MuddyMix. Because standard audio metrics reward technical accuracy over creative choices, the researchers instead borrowed evaluation methods from generative modeling.

In their own experiments, the researchers report that SSE beat the baseline systems they tested it against on both controllability and remix quality. That combination - steering a remix with plain language and video context together, rather than a mixing board - is the part worth watching. Most consumer audio-cleanup tools handle noise removal alone, not full remixing with source separation and reverb control in one text-guided pass.

It's a preprint, not a shipped product, and the wins are measured on the team's own new dataset against baselines they chose. Whether SSE holds up outside the lab, or against tools professional audio editors already trust, is still an open question.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →