Researchers have built a vision-and-language navigation system that knows when to stop and ask a human for directions.
The team's framework, called AVERT-VLN, adds a separate monitoring model that watches a navigation agent's progress against its instructions, comparing visual history and current observations to flag when the robot has wandered off course. To train that monitor, the researchers built the LOSTNAV dataset: 20,000 counterfactual trajectories engineered to go wrong, plus rule-based labels marking the deviation. The monitor is fine-tuned first on 40,000 normal trajectories, then on a mix of normal and risky ones, so it learns what a mistake actually looks like. When it issues a LOST verdict, the system pauses and waits for a human to correct course, then feeds that correction back into training through what the researchers call trajectory-anchored preference learning.
This matters because most vision-and-language navigation research chases full autonomy, treating any request for help as a failure mode. AVERT-VLN inverts that assumption: with human-assisted recovery, the full system hits success rates of 76.2% and 66.3% on the R2R-CE and RxR-CE val-unseen benchmarks, and the same monitor-and-recovery setup improved results across all three navigation architectures the team tested it on.
It is still a research prototype measured on curated benchmarks, not a robot loose in a real warehouse, so treat those success-rate numbers as a ceiling, not a guarantee.