[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-a-pushing-benchmark-exposes-a-blind-spot-in-ai-planning":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},9931,"a-pushing-benchmark-exposes-a-blind-spot-in-ai-planning","A Pushing Benchmark Exposes a Blind Spot in AI Planning","A new benchmark called SLIM shows popular world models barely plan pushing tasks until one fix repairs their blind, action-insensitive latent space.","Researchers built a benchmark that breaks a popular class of AI planning models, then found a fix that mostly repairs it.\n\nThe benchmark, called SLIM, asks a simulated robot arm to push small objects around a tabletop toward goals given either as an image or as a sentence. A world model called LeWM, which had aced a similar benchmark called PushT, solved under 1% of SLIM's trials, while a scripted controller with direct access to the simulator's internal state solved every tier. Diagnostic probes traced the failure to the model's encoder: its internal representation barely changed when the robot acted, and researchers couldn't even decode the pusher's or the objects' positions from it. Adding one extra training signal, a loss that forces the encoder to predict actions from pairs of latent states, fixed the probes and pushed success from 0.3% to 35%.\n\nThat's a quiet admission that a chunk of recent world model research may work by accident, succeeding only when enough of the frame happens to move for the model's shortcuts to still function. The paper's action-sensitivity probe, which needs no simulator access, offers a cheap pre-flight check before anyone trusts a world model's plans. That matters anywhere a company proposes letting a model simulate outcomes before a robot, or an agent, acts in the real world.\n\nOn the repaired model, a language-goal head reached 84% success on navigation tasks, versus 100% for the version given the goal as an image. Precision pushing is harder: a single sentence describing a goal rarely finishes a push, and even breaking the task into a sequence of stage-by-stage sentences only lifts success on the medium and hard pushing tiers from 4% to 25%. Useful, but a reminder that telling a robot what you want is still a weaker interface than showing it.","[\"ai\",\"robotics\",\"world-models\",\"benchmarks\"]","2026-10-05T04:00:00.000Z","2026-10-05T13:53:20.484Z","2026-10-05T13:53:26.765Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The closing line attributes the 0.25 success rate on the hardest pushing tier to 'a single sentence describing a goal,' but the source says a single sentence 'rarely completes a push' — the 0.25 figure applies to the multi-step sequence of stage sentences, not a single sentence; fix this misattribution so the single-sentence and multi-step results aren't conflated.","resolved","ai",[30,32,33,34],"robotics","world-models","benchmarks",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.03137",0,{"sections":41},[42,45,49,54,59,64,68,73,77,81,86,91,96,101],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",6168,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",859,{"name":50,"slug":51,"count":52,"latest_published_at":53},"Policy","policy",444,"2026-10-03T15:02:01.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",323,"2026-10-04T13:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",204,"2026-10-03T14:50:50.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":18},"Science","science",177,{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",158,"2026-10-03T03:21:12.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":18},"Dev Tools","dev-tools",97,{"name":78,"slug":79,"count":76,"latest_published_at":80},"Software","software","2026-10-04T10:00:00.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Startups","startups",92,"2026-10-04T14:36:25.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"General","general",51,"2026-10-05T02:35:01.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"Reviews","reviews",32,"2026-10-02T18:00:00.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]