A new AI editing method can paste an object from one video into another moving clip without retraining a model or tracking anything by hand.
MoCA-Video is a training-free framework built on an existing, frozen video-diffusion model - the AI doesn't need extra training data to do the edit. At specific steps late in the model's internal denoising process, it uses the model's own attention patterns to find a target object in the footage, then injects the appearance of a separate reference object into that spot, timed for the moment when the object's shape has formed but its surface look can still change. A momentum-based correction step carries that injected look across subsequent frames so the swapped-in object doesn't flicker or drift. On an Nvidia A100 GPU, rendering takes about 3.2 seconds per frame, not counting setup.
Most video-editing AI needs either a model trained specifically for that edit or a separate tool to locate objects in the frame first. MoCA-Video skips both, which could make this kind of object-insertion edit cheaper to build into other software. The strongest results, though, rest on a metric the same researchers created: CASS, a CLIP-based score that measures how far an edited video's imagery shifts toward the reference object and away from the original prompt. MoCA-Video reportedly scores highest on CASS and on ImageReward, a separate score estimating how appealing an image looks to people, but the paper does not publish the actual numbers or margins. Two other metrics tracking frame-to-frame smoothness (LPIPS-T) and overall video quality (FVD) reportedly show trade-offs rather than a clean win.
A three-second-per-frame render time is fine for a research paper. Real-time video editing this is not - yet.