A single AI system can now explain a factory defect, outline exactly where it sits on a product, and generate a realistic edit showing it fixed, all in one pass.
The system, called IAD-Unify, links a multimodal language model with a dense visual detector and a diffusion-based image editor. A shared visual pathway built on DINOv2 scans multiple reference images and compresses what it finds into a compact set of evidence tokens. Those tokens feed into a Qwen3.5 language model, which answers questions about the defect and converts the same evidence into a pixel-level segmentation mask. A separate interface then passes the language model's understanding to a Stable Diffusion-based editor, which uses the original image and a mask to generate a controlled edit of the defect.
Industrial quality control has typically needed one model to describe a flaw, another to locate it, and a third to simulate a repair, each built and maintained separately. IAD-Unify's developers report it reaches 73.02 percent on the MMAD Macro-7 benchmark and posts the highest published segmentation accuracy average among single-reference methods, at 57.10 percent, suggesting one shared model can match specialized tools without routing every task through the same bottleneck.
It is a benchmark result from a research paper, not a factory-floor product, but it points toward inspection software that explains itself instead of just flashing a red light.