A new benchmark says large language models are still bad chemists when it comes to picking reaction conditions.
Researchers built RxnOptBench, a test drawn from real condition-screening tables in organic-chemistry papers published in 2025. Models are asked to choose the catalyst, ligand, solvent, temperature, and other variables that maximize yield and selectivity, graded against a continuous score combining reported yield with enantiomeric excess, diastereomeric ratio, and regioisomeric ratio. The benchmark also runs each model with and without access to prior literature precedents, which separates genuine reasoning from models simply regurgitating memorized training data. Across nine frontier LLMs and three chemistry-specific models, the results were underwhelming.
The notable finding is that the chemistry-specialized models, built specifically for this kind of work, performed no better than random guessing on multi-variable selection. Meanwhile general-purpose open-weight models have nearly closed the gap with proprietary frontier systems, suggesting raw scale and broad training data matter more here than narrow domain fine-tuning.
It is a reminder that an LLM fluent in chemistry jargon is not the same as one that can run a lab bench - those two skills apparently do not transfer.