[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-claude-opus-5-favors-biology-over-evidence-in-study":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8880,"claude-opus-5-favors-biology-over-evidence-in-study","Claude Opus 5 Favors Biology Over Evidence in Study","A new study finds Claude Opus 5 often picks the biologically expected answer over the better supported one, while GPT-5.6 Sol barely budges.","Claude Opus 5 will quietly swap the better supported answer for the one that sounds more biologically plausible, according to a new study on how language models judge conflicting scientific claims.\n\nResearchers built a set of constraints that could not all be true at once, then translated them into lab-report-style prose where one answer satisfied more of the constraints and a different answer better matched what biologists would expect to see. Stated as plain logic, both models found the best supported answer reliably: GPT-5.6 Sol got it right 90% of the time, Claude Opus 5 96%. Once the same constraints were rewritten as scientific narrative, Claude Opus 5's accuracy on the evidence-based answer dropped to 27%, while GPT-5.6 Sol's performance barely moved. Removing the biological framing brought Claude Opus 5 back to 79%, and adding an explicit formalization request with a cue about the paired study design pushed it to 92%.\n\nThe gap matters because the failure is not about math. Both models can solve the constraint puzzle when it is presented as a puzzle. The weak point is earlier: deciding whether two findings are even measuring the same thing, and letting assumed biology quietly answer that question instead of the data.\n\nThat is a specific, fixable prompting problem, not evidence that one model reasons worse than the other. It is also a caution for anyone wiring these models into literature-review or hypothesis-checking tools: the same framing tricks that mislead a rushed human reviewer can mislead the model reading on their behalf.","[\"ai\",\"llm-evaluation\",\"ai-research\",\"anthropic\"]","2026-10-01T04:00:00.000Z","2026-10-01T09:02:01.446Z","2026-10-01T09:02:06.712Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The headline says 'AI Models' (plural) favor expectations over evidence, but the body shows only Claude Opus 5 exhibited this bias while GPT-5.6 Sol's performance barely moved — narrow the headline to reflect that the effect was specific to one model, matching the dek's accuracy.","resolved","ai",[30,32,33,34],"llm-evaluation","ai-research","anthropic",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.38621",0,{"sections":41},[42,45,50,55,60,65,70,75,80,84,89,94,99,104],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5351,{"name":46,"slug":47,"count":48,"latest_published_at":49},"Security","security",801,"2026-09-30T22:18:23.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",429,"2026-10-01T02:26:17.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Science","science",157,"2026-09-30T15:00:56.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":81,"slug":82,"count":78,"latest_published_at":83},"Software","software","2026-09-30T21:41:11.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":105,"slug":106,"count":107,"latest_published_at":108},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]