[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-graph-reasoning-method-lifts-llm-accuracy-166-on-benchmark":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},10666,"new-graph-reasoning-method-lifts-llm-accuracy-166-on-benchmark","New Graph Reasoning Method Lifts LLM Accuracy 16.6% on Benchmark","Researchers built a graph-reasoning method that looks ahead before pruning evidence, boosting a leading LLM benchmark's accuracy by over 16 percent.","A new evidence-retrieval method lets AI models think several steps ahead before deciding which facts to trust, and it beats prior approaches on a standard benchmark by double digits.\n\nResearchers built a system called Foresight-over-Graph (FoG) that helps large language models search knowledge graphs for question answering. Unlike existing methods that prune weak-looking evidence hop by hop as they go, FoG builds a question-relevant subgraph and uses feedback from deeper exploration to decide which paths were actually worth keeping. On the CWQ benchmark, a standard test for complex, multi-hop questions, FoG improved the Hit rate - how often the right answer shows up among the retrieved evidence - by 16.58%, while making fewer LLM calls and using fewer tokens than competing methods. The code is public on GitHub.\n\nKnowledge graphs are one of the more credible pitches for keeping LLMs honest on fact-heavy questions, since a graph is structured and checkable in a way a model's internal weights are not. The problem with earlier graph-search methods was that greedy, hop-by-hop pruning threw away branches that only turned out to matter several steps later, with no way to recover them. Fixing that ordering problem, rather than adding more training data, is what moved the number.\n\nA hit-rate gain on one benchmark is not proof that graph-grounded LLMs have solved hallucination - it is evidence that smarter evidence-retrieval helps on this particular, multi-hop-heavy test, and that is a narrower claim worth keeping straight.","[\"ai\",\"knowledge-graphs\",\"llms\",\"benchmarks\"]","2026-10-07T04:00:00.000Z","2026-10-09T02:39:11.854Z","2026-10-09T02:39:14.789Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The headline claims FoG 'cuts AI hallucinations,' but the paper only reports a Hit-rate\u002Faccuracy improvement on the CWQ benchmark — rewrite the headline and framing to state the actual benchmark result rather than implying a hallucination metric that wasn't measured.","resolved","ai",[30,32,33,34],"knowledge-graphs","llms","benchmarks",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.08388",0,{"sections":41},[42,46,51,56,61,66,70,75,80,84,89,94,99,104],{"name":43,"slug":30,"count":44,"latest_published_at":45},"AI",6506,"2026-10-07T18:45:00.000Z",{"name":47,"slug":48,"count":49,"latest_published_at":50},"Security","security",911,"2026-10-07T19:53:42.000Z",{"name":52,"slug":53,"count":54,"latest_published_at":55},"Policy","policy",474,"2026-10-07T18:23:21.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Deals","deals",453,"2026-10-07T23:58:31.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":65},"Hardware","hardware",222,"2026-10-07T21:19:54.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":18},"Science","science",187,{"name":71,"slug":72,"count":73,"latest_published_at":74},"Consumer Tech","consumer-tech",174,"2026-10-07T17:41:41.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Software","software",113,"2026-10-07T18:10:00.000Z",{"name":81,"slug":82,"count":78,"latest_published_at":83},"Startups","startups","2026-10-07T23:36:57.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Dev Tools","dev-tools",105,"2026-10-07T16:59:11.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"General","general",61,"2026-10-07T22:00:24.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"Gaming","gaming",56,"2026-10-07T12:00:00.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"Reviews","reviews",33,"2026-10-05T11:57:17.000Z",{"name":105,"slug":106,"count":107,"latest_published_at":108},"How-To","how-to",8,"2026-10-05T09:00:00.000Z"]