A new robot navigation system finishes a walking route as reliably as a human guide, and does it in well under half the time.
Researchers describe NavGPT-3, a runtime that splits navigation into separate reasoning, acting, and monitoring threads running on top of a language model and a vision-language-action policy called NavGPT VLA. NavGPT VLA alone, trained on 19.28 million examples, hits 74.51 percent success on the R2R-CE benchmark and leads RxR-CE at 78.19 percent. Layered with the full reasoning harness, NavGPT-3 pushes R2R-CE to a new high of 81.51 percent and, on RxR-CE, matches human performance: 90.43 percent success versus 90.4 percent for human followers, with similar path accuracy. It does that in 1 minute 22 seconds per episode, against roughly 3 minutes for a person - about 2.2 times faster.
The real story is the plumbing, not the score. By running reasoning and low-level control as separate, interruptible threads, the system lets slow, deliberate language-model thinking hand off to a fast action policy that reacts in 0.5 to 1 second, instead of the 3 to 19 seconds a pure reasoning loop needs per decision. That is the actual bottleneck holding back language-model-driven robots: not intelligence, but how fast that intelligence converts into motion.
The numbers come from simulated vision-and-language navigation benchmarks, not a robot wandering a real hallway, so treat human-level as a benchmark claim until it survives contact with an actual building.