AI researchers have found a fix for one of motion-customization's most persistent headaches: generated videos quietly copying the look of whatever reference clip they borrowed motion from.
The problem, known as content leakage, happens when a model trained to copy motion from a reference video also picks up that video's appearance - skin tone, outfit, background, whatever. The researchers trace this to how these systems are usually trained: treating the task as direct regression against the reference clip pushes the whole generative process to collapse toward that reference, dragging its look along with its movement. Their fix, called Control-based Motion Customization (CMC), reframes training using stochastic optimal control, a framework for steering a system's behavior without forcing it to mimic one fixed trajectory. That lets the model learn the target motion while leaving appearance to the text prompt, the way the base model intended. They also added a timestep-adaptive cost that only polices motion during the early stages of generation, which they say speeds up training by 2.5 times.
This matters because content leakage is the practical ceiling on tools that apply one video's movement to a different subject, from dance-transfer effects to synthetic stunt work. Until now, reducing leakage usually meant sacrificing how closely the output matched the reference motion. CMC claims to hold onto motion fidelity and the base model's output diversity at the same time, which is the harder trade-off to solve.
Still, "competitive motion fidelity" is the paper's own framing, not an independent benchmark result, and this is research code, not a shipped feature. Whether it holds up against messy, real-world reference footage is the question worth watching.