AI/ ai-video · generative-ai · computer-vision · diffusion-models

PISCO Lets Editors Insert Video Objects With Sparse Keyframes

A research model inserts objects into existing video footage from just a few keyframes, swapping prompt guesswork for precise, targeted control.

Researchers have built an AI model that drops a person or object into existing video, then figures out on its own how it should move and interact with the scene.

The system, called PISCO, is described in a new paper. Instead of demanding a fully-written prompt and endless regeneration, PISCO takes a small set of keyframes - as few as one - and propagates the inserted object's appearance, motion, and physical interactions across the rest of the clip. The team built new conditioning methods, including what they call Variable-Information Guidance and Distribution-Preserving Temporal Masking, to stop the model from falling apart when given so little input to work with. They also released PISCO-Bench, a benchmark with verified object annotations and matching clean-background footage, to measure how well insertion tools actually preserve a scene's original dynamics.

This matters because most AI video-editing tools still treat object insertion as a prompt-engineering puzzle - describe what you want and hope the model cooperates. PISCO's bet is that a handful of concrete keyframes is a more reliable lever than more words, and the paper's claim of steady, predictable improvement as more keyframes are added is the opposite of how most diffusion models tend to behave under sparse guidance.

It is a narrow fix, but the problem it targets - precise, repeatable video edits - is exactly what has kept generative video out of real production pipelines so far.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →