AI/ ai · llm-evaluation · ai-research · anthropic

Claude Opus 5 Favors Biology Over Evidence in Study

A new study finds Claude Opus 5 often picks the biologically expected answer over the better supported one, while GPT-5.6 Sol barely budges.

Claude Opus 5 will quietly swap the better supported answer for the one that sounds more biologically plausible, according to a new study on how language models judge conflicting scientific claims.

Researchers built a set of constraints that could not all be true at once, then translated them into lab-report-style prose where one answer satisfied more of the constraints and a different answer better matched what biologists would expect to see. Stated as plain logic, both models found the best supported answer reliably: GPT-5.6 Sol got it right 90% of the time, Claude Opus 5 96%. Once the same constraints were rewritten as scientific narrative, Claude Opus 5's accuracy on the evidence-based answer dropped to 27%, while GPT-5.6 Sol's performance barely moved. Removing the biological framing brought Claude Opus 5 back to 79%, and adding an explicit formalization request with a cue about the paired study design pushed it to 92%.

The gap matters because the failure is not about math. Both models can solve the constraint puzzle when it is presented as a puzzle. The weak point is earlier: deciding whether two findings are even measuring the same thing, and letting assumed biology quietly answer that question instead of the data.

That is a specific, fixable prompting problem, not evidence that one model reasons worse than the other. It is also a caution for anyone wiring these models into literature-review or hypothesis-checking tools: the same framing tricks that mislead a rushed human reviewer can mislead the model reading on their behalf.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →