AI/ ai · llm-evaluation · benchmarks · biomedical-ai

New Benchmarks Expose Gaps in LLM Scientific Reasoning

A new ontology-grounded benchmark pipeline shows top LLMs scoring just 41 to 77 percent on verified science reasoning questions.

A new testing pipeline checks whether AI models actually reason through science facts, or just guess well.

Researchers built an automated system that generates multiple-choice benchmarks straight from OWL 2 ontologies, the structured knowledge bases used in fields like biomedicine and materials science. Correct answers come directly from the ontology's own class definitions. Wrong answers are created by distorting those definitions in ways a formal logic reasoner can prove are false, not just plausible-sounding guesses. The team ran the pipeline on three ontologies: a small pizza-topping taxonomy, the materials-science ontology PMDco, and DOID, a large biomedical disease ontology, generating 112, 2,491, and 15,216 questions respectively. Six LLMs tested with no special training scored between 41.1% and 76.8% accuracy, well above the 25% random-guess baseline, but nowhere near solved.

That gap matters because LLMs already answer biomedical and materials-science questions in real tools, where a wrong answer that sounds confident is worse than an obvious one. The method also fixes a circular problem: grading an LLM's reasoning with another LLM risks just measuring shared blind spots, not actual logic. Here, every wrong answer is formally verified false by a reasoner before a model ever sees it.

A benchmark that proves its own trick questions are wrong is a refreshing change from ones assembled by committee, or quietly generated by another chatbot.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →