AI/ ai · image-to-video · video-generation · research

New Method Reduces Static Motion in AI Video Generation

A new training-free technique called DyMoS tweaks attention inside image-to-video models to reduce the stiff, static motion that has dogged the format.

A new paper claims to loosen up one of AI video generation's most annoying tics: footage that barely moves.

Researchers behind a preprint called DyMoS (Dynamic Motion Slider) say they found why image-to-video models produce such static clips: frames after the first one pay too much self-attention to the reference image's tokens, which over-propagates the reference forward through time and suppresses motion between frames. Their fix does not retrain the model or alter the input image. It rebalances how much attention later frames pay to the reference frame during the early denoising steps, controlled by a single adjustable parameter. The method is described as model-agnostic and training-free, and the authors say it worked across several state-of-the-art image-to-video backbones in their tests, improving motion dynamics while keeping visual quality and fidelity to the source image.

Image-to-video models have lagged text-to-video ones specifically because they cling too hard to the starting image, producing clips that look more like slideshows with added grain. A training-free tweak that reportedly works across multiple backbones is notable because most workarounds to date have required retraining or accepted a quality tradeoff. If DyMoS holds up outside the paper's own benchmarks, it is a cheap lever any lab could bolt onto an existing pipeline.

The claims come from the authors' own preprint, measured against their own benchmarks. Improving motion dynamics, per the abstract, is not the same as producing genuinely dynamic video, and it remains to be seen whether one adjustable slider closes the gap with text-to-video output.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →