A fix for a subtle flaw in AI video generators just showed up in a new paper — and it comes straight from the playbook that scaled up chatbots.
Researchers propose SplitMoE, a mixture-of-experts architecture built specifically for video diffusion models. Standard MoE systems route each token to an expert independently and push for roughly equal usage across the expert pool, a technique borrowed wholesale from language models. The paper argues that approach backfires on video, where neighboring frames are heavily redundant and some visual concepts show up far more often than others. Forcing uniform routing scatters related patches across unrelated experts, the authors say, which fragments routing and distorts the generated video.
Mixture-of-experts has become a standard way to scale large models without scaling compute costs in lockstep, and video teams have been borrowing the same trick without adapting it. SplitMoE splits the expert pool into two roles instead: semantic experts that handle high-level structure, and generic experts that mop up leftover visual detail, guided by prototype-based routing rather than a hard balancing rule. At an equal activated-parameter budget, the authors report faster convergence, more coherent routing, and better video quality than conventional load-balanced MoE setups.
It's one arXiv paper, not a shipped model, so the benchmark wins are a promising first read, not a settled result. The real test is whether this holds up at the scale needed for the video world models everyone is racing to build.