A new benchmark hands a robot a deck of cards, and it mostly folds.
Researchers behind DexHoldem built a real-world test rig where a ShadowHand robotic hand performs Texas Hold'em tasks like dealing and chip handling, backed by 1,470 teleoperated demonstrations across 14 manipulation primitives. On raw execution, the best policy, called pi-0.5, completed 61.2% of its assigned moves. On a stricter measure, scene-preserving success, meaning the table was left usable for the next step, pi-0.5 and an earlier version called pi-0 tied at 47.5%. The benchmark also scores whether an AI system can read the table and reconstruct the full game state, a separate perception test from the physical manipulation one.
On that perception test, the paper credits a model it calls "Opus 5.5" with the top score, 49.1% on strict accuracy and 80.6% on partial field-by-field accuracy. That name doesn't match any Anthropic model we can confirm has actually shipped, so the claim is worth flagging rather than repeating as settled fact. What's not in question is the harder result: when researchers ran the full agent-plus-robot loop across 33 hands, only 12.1% finished without a human stepping in, even with retries patching some failed moves.
Isolated skill tests are getting easier for robots to pass. Running an entire task end to end, where perception, decision-making, and dexterity all have to work together without a human bailout, is still the part nobody has solved.