AI/ reinforcement-learning · diffusion-models · robotics · arxiv

New Framework Merges Diffusion Models With Entropy RL

A revised arXiv paper argues diffusion-based policies beat standard entropy RL at finding varied, successful robot-control strategies.

A reinforcement-learning paper quietly revised this week argues diffusion models can replace the clumsy Gaussian policies most robots rely on.

Researchers introduce Diffusion-Augmented Markov Decision Processes, or DA-MDPs, which treat each step of a reverse-diffusion process as its own small reinforcement-learning decision, while only the final cleaned-up action actually gets executed by the robot. The method derives a tractable upper bound on the gap between the policy and its target, using the data-processing inequality, and splits that bound across each denoising step. The team plugged the framework into three existing algorithms, PPO, REPPO, and a maximum-entropy version of WPO, then tested it on two simulated robot-manipulation tasks, StackCube and PushT. DA-MDP policies hit higher success rates and found a wider range of distinct successful strategies than a standard Gaussian maximum-entropy baseline.

That variety matters more than the raw success numbers. Most reinforcement-learning policies converge on one way to solve a task and get brittle when conditions shift; a policy that keeps several working strategies in reserve is more useful once it leaves simulation. The paper also reports the approach trains with less memory overhead and still works when actions are grouped into chunks, a detail that matters for anyone trying to run this on real hardware rather than a cluster.

Worth noting: arXiv's own numbering, 2512.02019, places the paper's original submission in December 2025, and this week's listing is a fourth revision, not a new result. The gains are also limited to two simulated pick-and-push tasks, so treat higher success rates as a lab finding, not a robotics breakthrough yet.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →