AI/ robotics · dataset · vision-language-action · bimanual-manipulation

New Dataset Helps Two-Armed Robots Narrate Their Own Steps

A new 40,543-episode dataset and matching model let bimanual robots predict their own next subtask, cutting failure rates on tricky tasks dramatically.

A new dataset teaches bimanual robots to narrate their own next move, and the results look nothing like marginal fine-tuning gains.

Researchers released FineART, a densely-annotated dataset for two-armed robot manipulation: 40,543 episodes, 1,718 hours, and 533,913 subtask labels spanning 151 tasks. They paired it with FineART-VLA, a vision-language-action policy trained to predict its own next subtask rather than just react to a single top-level instruction. Mid-training the model this way pushed success on a spatial disambiguation task from 32.0% to 100.0%, and step-by-step human subtask guidance lifted an unseen long-horizon task from 16.0% to 76.0%. The dataset, model weights, and training code are all open-sourced.

Most bimanual datasets label an entire episode with one instruction and leave the robot to infer every intermediate step, which is a big reason multi-step manipulation still trips up current systems. Because FineART-VLA already reasons in subtasks, fine-tuning it for a new robot took one-tenth the data required by baselines without that mid-training, and it generalized zero-shot to tasks it had never seen on that new hardware. That kind of transfer is what actually determines whether a robotics approach scales past a single lab's arm.

These are still curated benchmark tasks, not a robot folding your actual laundry - but if subtask prediction generalizes as well outside these 151 tasks as the numbers suggest, it's a real methodological gain, not just a bigger pile of training data.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →