[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-graph-world-models-cut-llm-robot-planning-failures-sharply":10,"sections":34},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":29,"feedback":33,"feedback_at":22,"cost_usd":33,"total_tokens":33},6840,"graph-world-models-cut-llm-robot-planning-failures-sharply","Graph World Models Cut LLM Robot Planning Failures Sharply","A new framework called GAVEL pairs LLMs with an explicit graph world model to catch and fix robot planning errors before they cause costly failures.","A new framework called GAVEL wraps LLMs in a verification layer that catches and repairs robot planning mistakes before a robot acts on them.\n\nGAVEL pairs an LLM planner with an explicit graph world model that tracks object relationships, action preconditions and effects, and probabilistic beliefs about object locations the robot hasn't observed yet. Before the robot executes an LLM-generated step, the graph model simulates the consequence, flags any rule violations, and fixes the ones it can solve on its own, kicking the error back to the LLM only when the fix requires actual reasoning. For tasks that bundle multiple subtasks, GAVEL also uses its beliefs about where objects probably are to resequence the remaining work and cut down on wasted searching. Tested on the BEHAVIOR-1K benchmark across 100 long single tasks and 500 multi-task instructions using Qwen3-8B, GAVEL raised single-task success from 41.2% to 91.8% and multi-task success from 19.9% to 92.6%, while belief-based resequencing trimmed travel distance by about 5.4% compared with a static ordering.\n\nLLMs are good at parsing instructions like putting a mug in the dishwasher and then wiping the counter, but bad at knowing the dishwasher door has to open first, or where the sponge actually is once it isn't where the model assumed. Those failures are exactly what block LLM-driven planning from shipping in real robots, where a bad guess means a dropped object, not just a bad chatbot answer. GAVEL's numbers suggest most of that failure isn't a reasoning problem at all, it's a bookkeeping problem, and a straightforward graph tracking preconditions and beliefs fixes the bulk of it without touching the LLM.\n\nBEHAVIOR-1K is a simulated benchmark, not a warehouse floor, and building the graph world model still requires someone to hand-encode preconditions and effects for every action the robot might take, the kind of labor-intensive scaffolding that tends to get glossed over in a results section.","[\"ai\",\"robotics\",\"llm-planning\",\"world-models\"]","2026-09-18T04:00:00.000Z","2026-09-18T19:01:59.805Z","2026-09-18T19:02:11.720Z","published",null,[],"ai",[24,26,27,28],"robotics","llm-planning","world-models",[30],{"name":31,"url":32},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.19315",0,{"sections":35},[36,39,43,48,53,57,61,66,70,75,80,85,90,95],{"name":37,"slug":24,"count":38,"latest_published_at":18},"AI",4031,{"name":40,"slug":41,"count":42,"latest_published_at":18},"Security","security",654,{"name":44,"slug":45,"count":46,"latest_published_at":47},"Policy","policy",338,"2026-09-11T04:00:00.000Z",{"name":49,"slug":50,"count":51,"latest_published_at":52},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":54,"slug":55,"count":56,"latest_published_at":18},"Hardware","hardware",155,{"name":58,"slug":59,"count":60,"latest_published_at":18},"Science","science",121,{"name":62,"slug":63,"count":64,"latest_published_at":65},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":18},"Dev Tools","dev-tools",78,{"name":71,"slug":72,"count":73,"latest_published_at":74},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]