Vision-language models (VLMs) are great at describing a photo but surprisingly bad at basic spatial questions, like which object is closer to the camera. A new paper proposes fixing that gap by training on images the AI generated for itself.
The pipeline, called VisionFoundry, starts with nothing but a task name - say, "viewpoint recognition." An LLM writes matching questions, answers, and prompts for a text-to-image model, which then generates the actual pictures. A multimodal filter checks the results and tosses out bad samples before they ever reach training. The team used this process to build VisionFoundry-10k, a synthetic dataset covering 10 distinct perception tasks, and fine-tuned three open-source VLMs on it. The gains were consistent: Qwen2.5-VL-3B-Instruct improved 6.7% on the MMVP-pair benchmark and 10.5% on CV-Bench-3D, and the approach also helped when used as reinforcement-learning data instead of plain fine-tuning.
This matters because VLMs have been scaling fast on language while still fumbling simple visual logic, mostly because natural photo datasets rarely come labeled with the spatial or viewpoint detail these tasks need. Generating that supervision synthetically, with no manual annotation, is a cheaper and more scalable fix than waiting for better-labeled real-world data to show up.
It is also the same trick that rescued LLM training once real text got scarce, now aimed at pixels - which is reassuring, but benchmark gains on lab backbones are not the same as proof this holds up in whatever model actually ships to your phone.