A new robot-training method pushes AI policies to attempt things they have never done, using little more than human thumbs-up-thumbs-down feedback.
Researchers built PrefPI, a framework that steers pretrained robot policies, including diffusion policies and the PI0.5 flow-matching vision-language-action model, toward behaviors outside their original training distribution. Instead of just reinforcing actions the policy already leans toward, PrefPI treats preferred trajectories as their own data distribution and compares that against the robot's broader behavior prior, using classifier-free guidance to amplify the difference. Repeating that comparison in a loop nudges the policy, step by step, toward behaviors it rarely or never produced on its own. On real hardware, the method raised a robot's object-transport height from 10.7 cm to 19.8 cm using only 150 preference-labeled trajectories.
That distinction matters. Most preference-learning methods, including typical reinforcement-learning-from-feedback setups, mainly sharpen behaviors a model already leans toward rather than teaching it something new. PrefPI's trick, treating preference as a density-ratio signal that can be amplified, points to a cheaper route to genuinely new robot skills without expensive reward engineering or mountains of demonstration data.
Still, nearly doubling a lift height is one result on one task and one robot. The paper is a preprint, not a peer-reviewed study, and it is a long way from a single metric improvement to robots handling messy, multi-step jobs.