A new system called VOMMI trains robot manipulation skills using ordinary phone video instead of teleoperation rigs or specialized motion-capture gear.
Researchers behind the VOMMI framework recorded demonstrations with two synchronized cameras, one on the body and one on the hand, to capture both how a person moves through a space and how they handle objects. A reconstruction step called R2-VO cleans up the drift that normally creeps into camera-based motion tracking, turning raw video into usable trajectory data without any robot in the loop. The team collected 500 demonstration trajectories per task, held out 75 for evaluation, and compared results against 200 demonstrations gathered the traditional way, directly on a robot. The resulting policy then feeds into a vision-language-action model through an adapter that folds in the video-derived motion signals.
Embodied AI's bottleneck has always been data: teleoperation rigs are expensive, slow to use, and don't scale past a lab. If phone video can substitute for robot-collected demonstrations without sacrificing accuracy, it changes who can contribute training data and how cheaply. The results back that up on paper: the video-trained policy beat a robot-trained one on base-velocity error by 18.2%, kept comparable end-effector accuracy, and lifted success rates 8.3 percentage points over the OpenPI 0.5 baseline across three real-robot tasks.
Promising, but this is one arXiv paper testing three tasks with a few hundred trajectories each, so treat the numbers as an early signal, not a verdict on where mobile-manipulation training is headed.