A new training method lets robots figure out, on their own, which demonstrations in a dataset are actually worth copying.
Researchers built SynIL, short for Synergy-based Imitation Learning, to fix a known problem in robot training: real-world demonstration datasets mix skilled and sloppy examples, and that noise drags down performance. Prior filtering methods require hand-picking expert reference data or writing task-specific rules, which does not scale. SynIL instead borrows an idea from neuroscience: skilled movement shows a low-dimensional, coordinated structure called motor synergy. The system measures how strongly a demonstration exhibits that synergy and turns the measurement into a reward signal for each step of a trajectory, via self-supervised regression, without any human labeling.
Tested on D4RL locomotion benchmarks and Robomimic manipulation datasets gathered from multiple human operators, the synergy-based rewards tracked closely with the datasets' actual ground-truth rewards. SynIL beat standard behavior cloning outright, and in sparse-reward human teleoperation cases, it even beat offline reinforcement learning trained on real environment rewards, despite never seeing those rewards itself, while matching it elsewhere. That matters because hand-engineering reward signals is usually the expensive part of robot training; automating a decent proxy could make it cheaper to use the messy, abundant data robots already generate, instead of curated demonstrations nobody can produce at scale.
The results come from benchmark suites, not warehouses or hospitals, so whether synergy-based scoring holds up on genuinely unpredictable human movement is still an open question.