[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-forcing-ai-to-show-its-causal-reasoning-boosts-accuracy":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8205,"forcing-ai-to-show-its-causal-reasoning-boosts-accuracy","Forcing AI to Show Its Causal Reasoning Boosts Accuracy","A two-turn prompting method that makes language models externalize a structured causal graph before answering lifts their accuracy by double digits.","A new prompting trick makes large language models spell out the causal graph before they answer, and the accuracy jump is in the double digits.\n\nThe underlying test, called Corr2Cause, asks whether a causal claim holds true across every possible directed graph that fits a given set of correlations and conditional independencies. Researchers found that when a model just reasons in free-form chain-of-thought, it tends to skip that hard graph-matching problem and fall back on shallow pattern matching instead. Their fix, called Structured Thinking, splits the task into two turns: the model first writes out a typed, schema-constrained summary of the causal graph, then answers using that graph as its stated basis. On the full Corr2Cause test set, this lifted Qwen3.5-27B's F1 score from 73.0 to 86.4 against a strong baseline that already includes explicit causal-discovery instructions, a 13.4-point jump that held up statistically; averaged across three seeds, the gain settled at 8.1 points.\n\nA control group that got the same instructions but wrote free-form prose instead of a structured graph reached only 67.6, showing the format constraint itself, not just extra reasoning steps, does the work. When researchers deliberately scrambled the graph the model produced, accuracy fell 12 points, indicating the model is actually using what it writes rather than treating it as decoration. The same pattern held on a second model, Qwen3.6-27B, and on GPT-5.4-mini, plus a paraphrased, out-of-distribution version of the test.\n\nChain-of-thought prompting has always promised that showing your work improves reasoning; this result suggests that promise only cashes out when the work is forced into a shape the model cannot fudge.","[\"ai\",\"causal-inference\",\"llm-reasoning\",\"benchmarks\"]","2026-09-28T04:00:00.000Z","2026-09-28T16:49:30.296Z","2026-09-28T16:49:36.932Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The source lists Qwen3.6-27B, Paraphrase-OOD, and GPT-5.4-mini as where the pattern holds, but Paraphrase-OOD is an OOD test condition, not a model — fix the 'three other models, including GPT-5.4-mini' claim so it doesn't misidentify a dataset\u002Feval condition as a model.","resolved","ai",[30,32,33,34],"causal-inference","llm-reasoning","benchmarks",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.31071",0,{"sections":41},[42,45,49,54,59,64,68,73,78,83,88,93,97,102],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",4844,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",762,{"name":50,"slug":51,"count":52,"latest_published_at":53},"Policy","policy",399,"2026-09-27T18:39:02.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",266,"2026-09-28T14:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",189,"2026-09-28T10:52:40.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":18},"Science","science",151,{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",135,"2026-09-26T14:30:00.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Dev Tools","dev-tools",84,"2026-09-26T04:20:58.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"Startups","startups",76,"2026-09-25T18:33:59.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":94,"slug":95,"count":91,"latest_published_at":96},"General","general","2026-09-26T17:02:42.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",30,"2026-09-24T20:07:31.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]