A new diffusion model claims to generate synthetic single-cell RNA data that holds up even when the underlying measurements are noisy or incomplete.
Researchers describe LapDDPM, a conditional graph diffusion probabilistic model built to synthesize single-cell RNA sequencing (scRNA-seq) data. It combines graph-based inductive biases with score-based generative modeling, then adds a spectral adversarial perturbation step that jitters graph edge weights along principal spectral modes during training. That perturbation acts as a distributionally robust optimization framework, forcing the model to stay accurate even when the input graph's structure gets noisy. The team also extended the approach to spatial transcriptomics and multi-modal data, and tested it against five datasets, PBMC3K, Dentate Gyrus, HLCA, Visium, and 10x Multiome, where it beat existing baselines on distribution matching, manifold preservation, and downstream utility.
Single-cell data is notoriously sparse, expensive to collect, and easy to corrupt with technical noise, which makes synthetic data a real shortcut for filling out rare cell types or testing downstream pipelines. A generative model explicitly built to resist structural noise, rather than just fit clean training data, addresses a failure mode that has quietly undermined a lot of computational biology tooling.
It's worth remembering this is a preprint replacement, benchmarked on the authors' own dataset picks; the real test is whether it holds up on cell types nobody's trained a generator on yet.