A new benchmark shows humanoid robots are much better at picking a tool than at finishing the job with it.
Researchers released HumanoidToolBench, an 18-task benchmark that tests humanoid robots across three scenarios, three difficulty levels, and two tool-set conditions. It ships with ToolBook, a dataset of 3,100 demonstrations collected both in simulation and on a real Unitree G1 robot. The team ran seven policies in simulation and three on the physical robot. Across the board, robots that correctly chose a tool still frequently failed to complete the task with it.
That gap matters because tool use is the skill that would let a humanoid actually substitute for a person in a warehouse or workshop, not just wave at a crowd. Focused tests on the GR00T N1.7 model found selection accuracy dropped sharply on tools the robot hadn't trained on, and the robot kept executing a task even after being given an unrelated instruction, evidence that it was pattern-matching, not reasoning about the job.
It's a useful corrective to the polished demo reels: picking up a screwdriver on camera and knowing what to do with it are two very different problems.