Researchers have taught a robot dog to switch its priorities on command, without retraining its brain.
A team describes PROMO, a reinforcement learning method that lets a single policy control a quadruped's locomotion based on a preference set at runtime rather than baked in during training. Normal RL controllers hardcode a fixed trade-off between things like tracking a command, staying stable, and conserving energy. PROMO instead treats that trade-off as a dial an operator can turn after deployment. In simulation, the team sampled 100 different preferences and found 67 produced genuinely distinct, non-dominated behaviors, with a 0.843 correlation between the stated preference and the resulting behavior. The policy also transferred zero-shot to a real Unitree Go2, where changing the preference alone cut energy use by up to 30.4%, tracking error by 38.7%, and peak body-tilt deviation by 59.0% compared to a balanced setting.
That matters because robot deployments rarely have one fixed job. A warehouse quadruped might need to prioritize battery life on a long patrol and precise tracking when threading a tight aisle. Today that usually means training or fine-tuning separate models for each scenario. PROMO suggests one policy could cover that whole range, with the trade-off set like a configuration option instead of a retraining job.
The catch is that this is one lab's simulation-to-one-robot result, not a deployed product, and 67 non-dominated behaviors out of 100 preferences measures diversity, not how well any single behavior holds up in messy real-world conditions. The code is open-source, so expect other labs to stress-test the claim before it becomes standard practice.