Researchers built a half-billion-parameter AI model that watches an intersection and switches the traffic lights itself, and it gets ambulances through faster than anything else tested.
The model, called VLALight, feeds camera footage from multiple directions straight into a single vision-language-action system, paired with text instructions that tell it which camera view corresponds to which traffic movement and signal phase. That single step replaces the usual pipeline, where separate vision and language models first describe the scene in text, then hand that description to another system to decide on a signal change. Researchers say each conversion step loses visual detail and adds delay. In testing, VLALight cut pooled emergency-vehicle waiting time by 21.1% compared with VLMLight, a cascaded system that uses that older multi-step approach, and it ran in real time on local hardware.
Traffic systems built by bolting a vision model onto a language model onto a decision system carry exactly the kind of layered complexity that adds latency a moving ambulance cannot afford. VLALight's bet is that a smaller, single model that skips the text-description middleman is both faster and more accurate at spotting an emergency vehicle and clearing its path. It also generalized to intersections and traffic patterns it had not seen during training, which matters more than a lab benchmark, since most real cities have intersections nobody trained a model on.
That 21.1% edge was measured against another AI system, not against the traffic engineers and preemption hardware already running real intersections, so how VLALight stacks up against systems cities actually use is still unclear.