A new training method lets robots learn complex manipulation skills almost entirely in simulation, then move onto real hardware with no hand-tuned rewards.
ExpertGen starts with a diffusion policy trained on imperfect demonstrations - either human teleoperation data or examples synthesized by large language models. Reinforcement learning then nudges that diffusion model's initial noise toward task success while keeping the underlying policy frozen, which keeps exploration inside safe, human-like motion patterns. That constraint lets the system learn from sparse rewards alone, skipping the reward engineering that usually trips up robotic RL. On industrial assembly tasks the approach hit a 90.5% success rate in simulation, and 85% on harder, long-horizon manipulation tasks, beating every baseline tested.
The real bottleneck in robot learning has always been data: teleoperated demonstrations are slow and expensive to collect at scale, and synthetic ones are usually low quality. ExpertGen's workaround - treating a frozen diffusion model as a steerable prior instead of retraining it from scratch - offers a cheaper route to expert-level policies that doesn't require pristine demonstrations to start from.
The team did distill these simulated experts into visuomotor policies and run them on real robots, but the paper stops short of reporting a real-world success rate. Until that number surfaces, the 90.5% and 85% figures are a strong simulation result, not proof the robots are ready for a warehouse floor.