A new robotics framework teaches robots to tell nearly identical instructions apart - and actually act on the difference.
Researchers built T3DP, a Text-to-3D Policy framework, to fix what they call unseen specification generalization - robots failing when instructions specify fine details, like an exact target position, how far to move something, or how open a drawer should be, that weren't covered by training demonstrations. Most text-to-3D policies compress an instruction and its demonstration into one global embedding, which blurs instructions that are similar but behaviorally distinct. T3DP instead keeps the token-level structure of both the language and the robot's movements, linking specific words to specific segments of behavior. That representation then conditions a point-cloud-based 3D diffusion policy without changing its underlying architecture.
Across the Meta-World, ManiSkill, and RoboTwin simulation benchmarks, T3DP beat global alignment by 11.0 to 14.2 points on held-out instructions across all 15 task families, and on real robots it raised average success from 47.5% to 65.0%. That's a meaningful gap: most instruction-following policies grasp the gist of a command but not its fine print, which works for picking up a cup but breaks down when the exact distance moved is what matters.
It's still a benchmark result, not a warehouse robot, but it points at a real gap between language fluency and physical precision that bigger models alone haven't closed.