A frozen drone-navigation model falls apart the moment people talk to it like humans, not like flight logs.
Researchers tested OpenFly, an aerial vision-and-language navigation agent, on two instruction styles: the detailed, trajectory-aligned commands it was trained on, and shorter, intent-driven phrasing closer to how people actually speak. Swapping in human-style instructions cut the navigator's success rate from 31.03% to 11.33%. A first attempt to close that gap, using a language model prompted with human writing examples to generate synthetic "Weak" commands, only nudged success back up to 15.27%. The researchers then built the Trajectory-Grounded Instruction Translator (TGIT), a front-end that keeps the navigator untouched and instead learns, from the navigator's own trajectory outcomes, how to rephrase weak commands into language the navigator actually follows.
TGIT raised success on weak commands to 37.93%, and transferred zero-shot to real human instructions, lifting success from 11.33% to 32.51%. It also improved results on held-out OpenFly data (4.95% to 20.79%) and produced gains on two other benchmarks, CityNav and AirVLN. None of that required retraining the underlying navigation model - just adding a translation layer in front of it.
It's a cheap fix for an expensive problem: the agent wasn't bad at flying, it was bad at listening.