A new technique called ARC makes today's robot models dramatically more capable at zero added cost in demonstrations or training scale.
The robotics field usually chases better performance with bigger models, more demonstrations, and more expensive training runs. Researchers behind ARC propose a cheaper fix: teach the robot to reason before it acts. Their recipe has three parts - reasoning traces that explain why a given action makes sense and what it should accomplish, an automatic labeling pipeline that generates those traces from existing demonstrations, and a tailored fine-tuning approach for adapting pretrained models to use them. They built one such trace set, ARC-Trace-DROID, entirely from the existing DROID dataset, without collecting a single new robot demonstration. Applied to models including pi_0.5 and Cosmos3-Nano-Policy, ARC set new state-of-the-art results on the RoboLab-120 and MolmoSpaces benchmarks, with gains of up to 50 percentage points on a reasoning-focused subset. On a real robot, it lifted pi_0.5's task success rate by 82.2 percentage points.
That's the real story here: capability gains that don't require the industry's usual scale-and-spend playbook. If reasoning traces can be mined from data robotics labs already have, smaller teams get a shot at state-of-the-art performance without foundation-model budgets.
Robotics research has a long history of benchmark numbers that don't survive contact with messier real-world conditions, so the real test is whether ARC's gains hold up outside RoboLab and MolmoSpaces.