AI/ robotics · reinforcement learning · ai agents · benchmarks

Robot Coding Agents Learn When to Trust Which Skill

A new coordinator learns from counterfactual outcomes to pick the right robot skill for the moment, boosting task success past prior code-as-policy systems.

Researchers have built a system that teaches robots which of their own skills to trust, task by task.

The approach, called RoboAware, sits on top of coding agents that already mix hand-written robot skills with frozen end-to-end policies. Instead of hand-tuning when to use which, the researchers trained a separate coordinator to pick the right policy family based on the robot's current state. They built a five-stage skill framework called P5, then used a technique called State-Locked Counterfactual Branching to replay the same moment and test every available skill family on it, something normal training never observes. Those replayed outcomes feed into Execution-Aware Learning, which pairs Monte Carlo tree search with Q-learning to turn the results into a model of which family works best in which situation.

On 100 tasks spanning three robotics benchmarks, RoboAware hit a 77.0% overall success rate, including 90.0% on RoboSuite and on RoboTwin's bimanual tasks, and 73.8% on the more varied LIBERO-Pro suite, beating existing code-as-policy and vision-language-action harness baselines. That matters because most robot systems fail not from bad skills but from bad judgment about which skill applies right now. Letting a model learn that judgment from counterfactual replays, rather than from someone's hand-written rules, is a cleaner way to scale.

It is a modest result for now, limited to 100 tasks in simulated benchmarks, but the counterfactual-branching trick is the kind of idea that tends to get borrowed.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →