An AI tool built to score how safe a street feels turns out to agree with people only weakly once it leaves the lab.
Researchers built a pedestrian routing system that estimates perceived safety from street-level photos. Rather than scoring images directly, it first generates a plain-language caption for each photo, then derives a risk score from structured features of that text, so every judgment stays inspectable instead of hiding inside a black box. The team tested nine captioning setups across five model families and found the caption-based scores matched a direct image-embedding baseline rather than trailing it. They then ran the pipeline across 654,115 images covering 36 electoral wards in Manchester and Huddersfield.
When checked against 3,669 ratings collected from real people in the field, the AI's agreement with human judgment was modest: a correlation of 0.262, well short of the 0.737 ceiling set by how much the human raters agreed with each other. A follow-up test made the pipeline 44% stronger on its lab benchmark and saw no measurable improvement in the field (r=0.250, p=0.84), evidence that benchmark gains here did not translate to real-world accuracy.
The system still reroutes walkers toward lower-risk streets, adding a median 12.78% more route length on trips of 3 to 6 km. That's a real behavioral change built on a signal that, by the researchers' own numbers, still falls well short of matching how people actually judge a street.