A new benchmark called SciR asks a blunt question: can large language models actually reason through science, or are they just pattern-matching text that looks scientific?
Researchers built SciR around three types of inference that show up constantly in science: deduction, induction, and causal abduction. Rather than relying on human-annotated scientific papers, which are expensive and hard to verify, or synthetic logic puzzles, which look nothing like real research writing, the team generated tasks from formal structures such as deduction trees, inductive rules, and causal graphs. Those structures were then rendered into realistic, multi-document scientific prose using domain-specific genres. The setup lets researchers independently control two separate difficulties: how hard it is to extract the relevant information from the text, and how hard the actual inference is once that information is found.
Six models were tested, and both difficulty axes hurt every one of them, with the effects compounding when stacked together. That is a more nuanced failure mode than most benchmarks capture. Reasoning models like DeepSeek-R1 outperformed plain instruct models specifically on the inference axis, while extraction difficulty turned out to be a separate weakness that existing science benchmarks rarely isolate. Even neurosymbolic pipelines, which hand off the actual logic to a verified solver, still stumbled when the surrounding text got harder to parse, meaning bad prose can sabotage even provably correct reasoning engines.
It is a useful corrective to "reasoning model" marketing, which rarely specifies what kind of reasoning improved, or how much of the reported difficulty was really just finding the right sentence.