Researchers have built a system that teaches robots which of their own skills to trust, task by task.
The approach, called RoboAware, sits on top of coding agents that already mix hand-written robot skills with frozen end-to-end policies. Instead of hand-tuning when to use which, the researchers trained a separate coordinator to pick the right policy family based on the robot's current state. They built a five-stage skill framework called P5, then used a technique called State-Locked Counterfactual Branching to replay the same moment and test every available skill family on it, something normal training never observes. Those replayed outcomes feed into Execution-Aware Learning, which pairs Monte Carlo tree search with Q-learning to turn the results into a model of which family works best in which situation.
On 100 tasks spanning three robotics benchmarks, RoboAware hit a 77.0% overall success rate, including 90.0% on RoboSuite and on RoboTwin's bimanual tasks, and 73.8% on the more varied LIBERO-Pro suite, beating existing code-as-policy and vision-language-action harness baselines. That matters because most robot systems fail not from bad skills but from bad judgment about which skill applies right now. Letting a model learn that judgment from counterfactual replays, rather than from someone's hand-written rules, is a cleaner way to scale.
It is a modest result for now, limited to 100 tasks in simulated benchmarks, but the counterfactual-branching trick is the kind of idea that tends to get borrowed.