A new dataset teaches bimanual robots to narrate their own next move, and the results look nothing like marginal fine-tuning gains.
Researchers released FineART, a densely-annotated dataset for two-armed robot manipulation: 40,543 episodes, 1,718 hours, and 533,913 subtask labels spanning 151 tasks. They paired it with FineART-VLA, a vision-language-action policy trained to predict its own next subtask rather than just react to a single top-level instruction. Mid-training the model this way pushed success on a spatial disambiguation task from 32.0% to 100.0%, and step-by-step human subtask guidance lifted an unseen long-horizon task from 16.0% to 76.0%. The dataset, model weights, and training code are all open-sourced.
Most bimanual datasets label an entire episode with one instruction and leave the robot to infer every intermediate step, which is a big reason multi-step manipulation still trips up current systems. Because FineART-VLA already reasons in subtasks, fine-tuning it for a new robot took one-tenth the data required by baselines without that mid-training, and it generalized zero-shot to tasks it had never seen on that new hardware. That kind of transfer is what actually determines whether a robotics approach scales past a single lab's arm.
These are still curated benchmark tasks, not a robot folding your actual laundry - but if subtask prediction generalizes as well outside these 151 tasks as the numbers suggest, it's a real methodological gain, not just a bigger pile of training data.