AI/ ai · research · datasets · audio-visual

Researchers Pair Audio and Video to Describe Urban Scenes

A new dataset combining audio and visual AI descriptions of city scenes pushes urban scene classification accuracy to 95.4%.

A new research dataset pairs sound and sight to describe city scenes, and the combination beats either sense alone.

The dataset, called AVSD-Scenes, contains 12,291 audio-visual scene descriptions built from the TAU Urban Audio-Visual Scenes dataset. Researchers first generated separate descriptions using two existing AI models: Qwen2-Audio-7B for audio and Qwen2.5-VL-7B for video. They then combined those modality-specific descriptions using three large language models - Qwen3-14B, Mistral-Small-3.2-24B-Instruct-2506, and Gemma-3-27B-it - to produce single multimodal descriptions. The team evaluated the results through semantic alignment checks, cross-modal retrieval, scene classification, an LLM-as-a-judge evaluation, and human review.

The merged descriptions reached 94.5% accuracy on urban scene classification, and stacking audio, visual, and text embeddings together pushed that to 95.4%. That gap over single-sense descriptions matters because it suggests the dataset captures real scene information rather than a labeling shortcut - accuracy held up even when scene labels were removed from the prompts used to generate the text.

It is a modest accuracy bump wrapped in a lot of model names, but for anyone building systems that need to describe a scene rather than just tag it, a usable dataset is worth more than another leaderboard entry.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →