A new machine learning model predicts how a molecule's ion drifts through gas in a mass spectrometer - a number called collision cross section, or CCS - more accurately than earlier tools, including ones built on physics instead of data.
CCS is one of the main numbers chemists use to identify an unknown compound: alongside mass, it tells you roughly how big and what shape the molecule's ion is. The new model, called GRACE, was built by adapting a pretrained 3D molecular geometry encoder and training it against a baseline that already accounts for the ion's chemical adduct - the extra atom or group that gives the molecule its charge. GRACE also feeds adduct identity into the model early, through a learned token and lightweight attention adapters, rather than bolting it on as an afterthought. Tested on more than 9,000 experimental molecule-adduct measurements, GRACE came out ahead of every other learned model across three difficulty levels, including a split specifically designed to test adducts the model had never seen, and it beat four previously published physics-based methods on a held-out set.
That adduct-sensitive split is the real test. Most CCS predictors treat the adduct as a minor detail, which means they tend to fall apart when a sample throws an ionization state they were not trained on - exactly the situation mass spec users hit constantly in real-world metabolomics and drug discovery work. By baking adduct identity into the model from the start, GRACE's authors are targeting that specific failure mode rather than just chasing a lower average error.
It is a solid, incremental result on a research benchmark, not a transformation of how mass spec works day to day - the training set is still a few thousand records, and the real proof will be whether GRACE holds up on compounds nobody curated in advance.