AI/ reinforcement-learning · robotics · ai-research · machine-learning

FERPO Algorithm Sidesteps a Known Flaw in Robot Control AI

A new reinforcement learning algorithm skips critic gradients to avoid unreliable updates, and claims faster actor training than REPPO.

A new reinforcement learning algorithm called FERPO trains robotic control systems without a shaky mathematical step that other top methods rely on.

Most state-of-the-art reinforcement learning methods for continuous control, think robotic arms or simulated hands, improve a policy by taking the gradient of a critic (a model trained to predict how much reward an action will earn) with respect to the action itself. That is essentially asking the critic which direction to nudge the controls. The catch: a critic trained to predict rewards accurately does not necessarily give accurate directional advice, so this step can produce unreliable updates. FERPO skips it entirely, instead calculating an ideal target distribution of actions from the critic's values, keeping that target close to the policy's recent behavior through entropy and KL-divergence regularization, and training the actor to match it via a statistical reweighting technique called self-normalized importance sampling.

This matters because continuous control, the category of RL that governs how robots and simulated agents move, is notoriously finicky to train, and small gradient errors compound over thousands of updates into visibly worse behavior. On the MuJoCo Playground and ManiSkill benchmarks, FERPO matched or beat existing methods on sample efficiency and updated its actor faster than a comparable algorithm called REPPO, hinting at a cheaper and more theoretically sound way to train these systems.

The gains, so far, exist only in simulation. Benchmarks are a long way from a warehouse robot or a factory arm, and that gap is where plenty of promising RL papers quietly stop mattering.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →