[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-training-method-lets-robots-learn-from-failed-attempts":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},8518,"new-training-method-lets-robots-learn-from-failed-attempts","New Training Method Lets Robots Learn From Failed Attempts","DEWO retrains a robot's internal world model from its own failed and successful attempts, lifting real-world task success by over 20 points in early tests.","A new training approach lets robots learn from their own screwups, not just from copying successful moves.\n\nResearchers built what they call World-Action Models, systems that generate robot actions while predicting how those actions will physically play out, and found that current post-deployment training mostly just tweaks behavior without sharpening those physical predictions. That is a problem for dexterous manipulation, where small execution errors compound fast and push a robot into situations its model never trained on. Their new method, Direct Experience World-Model Optimization (DEWO), identifies the moment an interaction is about to succeed or fail and learns from both outcomes, while a separate module monitors task progress from video and adds extra guidance when the robot stalls. Across five simulated DexJoCo tasks, DEWO improved success rates for all three model variants tested, and on real Wuji and Sharpa robots running four tasks, two rounds of this learning raised success from 51.0% to 71.7% in grid cells with at least one prior success.\n\nThat distinction, tuning the world model itself rather than just the policy on top of it, matters because it targets the actual bottleneck in robot learning: bad predictions, not just bad choices. A 20.7 percentage point jump on a real robot, not just in simulation, is the kind of result that's easy to overstate but hard to dismiss.\n\nStandard post-deployment fine-tuning for robots usually just adjusts which actions a policy picks, on the assumption that a model's internal sense of physics carries over fine from pretraining. DEWO's bet is that this assumption breaks down exactly where robots need help most: novel, borderline failure states outside the training distribution. But the real-world evidence here comes from two robot platforms, four tasks, and 3x3 grid evaluations, not the sprawling variety of a warehouse or a home. Confirming this generalizes will mean testing across many more tasks and environments, running longer deployment windows, and seeing the results replicated by teams outside the original group, before anyone should treat grid-cell success rates as a preview of what ships.","[\"robotics\",\"world models\",\"machine learning\",\"arxiv\"]","2026-09-30T04:00:00.000Z","2026-09-30T08:36:51.685Z","2026-09-30T08:36:55.000Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The closing paragraph is a single caveat sentence with nothing else — expand the ending with additional context (e.g., what would need to happen next to validate this outside the lab, or how DEWO's approach compares to standard post-deployment fine-tuning) rather than trailing off on the benchmark-skepticism line alone.","resolved","ai",[32,33,34,35],"robotics","world models","machine learning","arxiv",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.37398",0,{"sections":42},[43,46,50,54,59,64,69,74,79,84,89,94,99,104],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",5028,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",780,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Policy","policy",417,{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":68},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":83},"Dev Tools","dev-tools",89,"2026-09-29T17:15:00.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":105,"slug":106,"count":107,"latest_published_at":108},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]