AI/ robotics · reinforcement-learning · quadrupedal-robots · ai

New RL Method Lets a Robot Dog Switch Priorities on the Fly

A new reinforcement learning policy lets a Unitree Go2 quadruped trade off speed, stability, and energy efficiency at runtime, not training time.

Researchers have taught a robot dog to switch its priorities on command, without retraining its brain.

A team describes PROMO, a reinforcement learning method that lets a single policy control a quadruped's locomotion based on a preference set at runtime rather than baked in during training. Normal RL controllers hardcode a fixed trade-off between things like tracking a command, staying stable, and conserving energy. PROMO instead treats that trade-off as a dial an operator can turn after deployment. In simulation, the team sampled 100 different preferences and found 67 produced genuinely distinct, non-dominated behaviors, with a 0.843 correlation between the stated preference and the resulting behavior. The policy also transferred zero-shot to a real Unitree Go2, where changing the preference alone cut energy use by up to 30.4%, tracking error by 38.7%, and peak body-tilt deviation by 59.0% compared to a balanced setting.

That matters because robot deployments rarely have one fixed job. A warehouse quadruped might need to prioritize battery life on a long patrol and precise tracking when threading a tight aisle. Today that usually means training or fine-tuning separate models for each scenario. PROMO suggests one policy could cover that whole range, with the trade-off set like a configuration option instead of a retraining job.

The catch is that this is one lab's simulation-to-one-robot result, not a deployed product, and 67 non-dominated behaviors out of 100 preferences measures diversity, not how well any single behavior holds up in messy real-world conditions. The code is open-source, so expect other labs to stress-test the claim before it becomes standard practice.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →