AI/ dataset-distillation · semantic-segmentation · diffusion-models · computer-vision

D3S2 Uses AI Image Generation to Shrink Training Data

A new diffusion-guided technique compresses semantic segmentation datasets to one percent of their size while still training usable models.

A new technique squeezes a full semantic segmentation dataset down to one percent of its size and still trains a usable model.

Researchers built D3S2, a two-stage method that first picks a balanced set of segmentation masks, prioritizing rare classes that normally get drowned out by common ones. It then feeds those masks into a pretrained layout-to-image diffusion model to generate matching photos. Two guidance signals, a pixel-alignment loss and a per-class feature-matching loss, steer the diffusion sampling so the generated images and labels line up and mimic the statistics real training data would have. Tested on the ADE20K and COCO-Stuff benchmarks with a Mask2Former model, the 1-percent-sized synthetic datasets hit 24.99% and 35.49% mIoU respectively, beating random selection of real images by 9.34 and 5.70 percentage points.

Dataset distillation has mostly focused on simple classification tasks, where each image needs only one label. Segmentation needs a label for every pixel, which makes class imbalance and spatial alignment much harder problems, and this work shows a diffusion model's generative flexibility can substitute for raw scale. If it holds up outside benchmarks, it points toward training segmentation models without hauling around terabytes of labeled imagery.

Beating random selection is a low bar, though. The real test is whether this beats other distillation baselines by enough to matter outside a lab.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →