A new research system lets video creators direct both camera movement and object motion in 3D space, starting from a single still image.
GenCine, built by researchers publishing on arXiv, turns one photo into an editable 3D scene scaffold. Artists draw a camera path and move parts of a subject using local 3D handles, which can act independently to approximate non-rigid motion without a physics engine. The system converts these controls into guidance maps that track each handle's position across frames in the same coordinate space as the background. A lightweight guidance branch and LoRA adapters trained on a pretrained Wan video model follow those maps to generate the final clip.
This matters because most controllable video tools today use flat 2D trajectories or drag signals, which get confused when the camera and the subject move at the same time. The same 2D line can mean different things in 3D, depending on where the camera sits. GenCine's approach keeps object motion anchored to the scene itself, so a tracked hand or car stays consistent even as the viewpoint shifts. That is a real gap in current AI video generation, not a cosmetic upgrade.
It is still a research prototype trained on synthetic and recovered real-world motion data, not a shipped editing tool, so the usual gap between an arXiv demo and something you can actually use in a video app still applies.