AI/ ai · science · genomics · machine-learning

CellMSA Borrows Protein Alignment Tricks for Cell Biology

Researchers unveil CellMSA, a model that borrows protein alignment tricks to read single-cell gene data across experiments and cell types.

A new single-cell AI model borrows a trick from protein science: look at relatives, not just the individual.

Researchers have released CellMSA, a representation-learning framework for single-cell transcriptomics data. Instead of encoding each cell in isolation, it retrieves related cells from other experimental batches and biologically similar cell types to build a context-dependent gene-pair representation before feeding it into the model. The approach mirrors multiple sequence alignment, a technique from protein modeling where sequences are compared against related ones to surface patterns. The team pretrained CellMSA on roughly 109 million cell observations, including 65.6 million primary ones, and reports it beats existing methods across multiple benchmarks. Code is available on GitHub under the PharMolix organization.

Single-cell sequencing data is notoriously messy - sparse, high-dimensional, and riddled with batch effects from different labs and instruments. Most prior foundation models just denoise within a single batch, discarding cross-batch and cross-cell-type signal that could sharpen gene-gene relationships. Borrowing context modeling from protein science, where it already works, is a sensible bet for making embeddings used in cell-type classification and perturbation prediction genuinely more useful.

Whether the gains hold outside the benchmark suite, and on cell types absent from that 109-million-cell training corpus, is the real test still to come.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →