[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-a-new-ai-model-skips-the-text-and-reasons-over-graphs-instead":10,"sections":50},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":39,"tags":40,"sources":45,"feedback":49,"feedback_at":22,"cost_usd":49,"total_tokens":49},6995,"a-new-ai-model-skips-the-text-and-reasons-over-graphs-instead","A New AI Model Skips the Text and Reasons Over Graphs Instead","A new NLI pipeline that never reads raw text trades a few points of accuracy for reasoning steps humans can actually audit.","Researchers built a natural-language-inference system that classifies sentence pairs without ever letting its classifier see the sentences.\n\nThe pipeline first breaks each sentence into atomic propositions, then converts those into ConceptNet triples using constrained decoding. Premise, hypothesis, and a retrieved ConceptNet subgraph are each turned into a graph, and only those three graphs get fed into a fine-tuned 0.8-billion-parameter model. On SNLI, the graph-only system hit 89.7% accuracy, just 1.9 points behind an identically trained text-based model. On ANLI, it matched published RoBERTa-large numbers on the harder rounds - 48.0% vs. 48.9% on R2, 44.9% vs. 44.4% on R3 - but fell 16 points behind on R1. Measured separately against the researchers' own text-based counterpart model, the system's overall ANLI gap came out to 9 to 14 points, a different baseline than the round-by-round RoBERTa-large comparison. Combining graphs with text pushed SNLI accuracy to 92.1%, ahead of either approach alone.\n\nThe pitch here isn't raw accuracy, it's a paper trail. Every decision the classifier makes can be traced back to specific extracted propositions and graph edges, instead of buried somewhere in a transformer's attention weights. That matters for any setting where someone needs to check why a model called two statements contradictory, not just accept the verdict.\n\nCall it the price of interpretability, which is what the researchers themselves call it: legible reasoning costs a few accuracy points on easy cases and a lot more on the adversarial ones. R1's 16-point shortfall is the tell - that round is designed to trip up simple pattern-matching, and it's exactly where compressing a sentence into propositions and graph triples throws away the nuance a full-text model would have caught.","[\"nli\",\"interpretability\",\"natural-language-processing\",\"conceptnet\"]","2026-09-18T04:00:00.000Z","2026-09-19T03:20:56.080Z","2026-09-19T03:21:07.987Z","published",null,[24,30,35],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Attribute the findings to their actual source — cite the arXiv paper (ID 2609.16814) and link it, since the draft currently attributes everything to an anonymous 'research team'\u002F'researchers' with no institution or publication reference.","resolved",{"id":31,"reviewer":32,"round":33,"reason":34,"status":29},"publisher-r2","publisher",2,"The stated ANLI results are internally inconsistent — matching RoBERTa-large on R2 and R3 (implying ~0 gap there) while trailing by 16 points on R1 cannot average out to an overall gap of 9 to 14 points.",{"id":36,"reviewer":32,"round":37,"reason":38,"status":29},"publisher-r3",3,"The ANLI numbers are internally inconsistent (R1 gap of 16 points alongside near-zero gaps on R2\u002FR3 cannot average to the claimed 9-14 point overall gap), which fails the fact-consistency check.","ai",[41,42,43,44],"nli","interpretability","natural-language-processing","conceptnet",[46],{"name":47,"url":48},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.16814",0,{"sections":51},[52,56,61,66,71,76,81,86,90,95,100,105,110,115],{"name":53,"slug":39,"count":54,"latest_published_at":55},"AI",4114,"2026-09-19T18:33:46.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Security","security",673,"2026-09-19T17:30:00.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":65},"Policy","policy",345,"2026-09-19T19:22:20.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Hardware","hardware",156,"2026-09-19T11:00:00.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Science","science",129,"2026-09-19T19:45:00.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":18},"Dev Tools","dev-tools",78,{"name":91,"slug":92,"count":93,"latest_published_at":94},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"General","general",42,"2026-09-18T22:35:10.000Z",{"name":111,"slug":112,"count":113,"latest_published_at":114},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":116,"slug":117,"count":118,"latest_published_at":119},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]