AI/ ai · benchmarks · chemistry · drug-discovery

Benchmark Finds AI Models Struggle to Read Chemical Patent Shorthand

A new diagnostic test shows vision-language models ace easy chemistry questions but falter badly once patent shorthand gets genuinely hard to parse.

A new benchmark shows that AI models still can't reliably read the fill-in-the-blank chemistry notation patent lawyers use to describe entire families of drugs.

Researchers built R-GroundBench, a diagnostic test drawn from real pharmaceutical patents, to probe how well AI models handle Markush structures - the placeholder notation (R1, R2, X, and so on) chemists use to describe whole families of related molecules in a single diagram. The benchmark includes multiple-choice questions at varying difficulty levels plus an open-ended generation task. Models did fine on easy questions, scoring over 90% accuracy. But accuracy dropped to 56-66% on harder questions designed to block shortcuts, and chemistry-specific models fared worse, managing only 25.7-46.2% despite being trained specifically on chemical data.

Markush structures are the backbone of how pharma companies protect intellectual property, covering thousands of possible compounds in a single patent claim. If AI tools are going to help with drug discovery or patent analysis, they need to parse this shorthand correctly, not just recognize familiar-looking diagrams. The open-ended generation task exposed an even bigger gap: exact-match accuracy fell below 20% for most models and below 8% when a model had to work from an image rather than text alone.

Specialized training didn't help much here - the chemistry-tuned models underperformed general ones on the hardest questions, a reminder that domain pretraining doesn't automatically buy domain reasoning.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →