[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-a-new-monitor-catches-bad-reasoning-in-robot-ai-sometimes":10,"sections":34},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":29,"feedback":33,"feedback_at":22,"cost_usd":33,"total_tokens":33},9649,"a-new-monitor-catches-bad-reasoning-in-robot-ai-sometimes","A New Monitor Catches Bad Reasoning in Robot AI, Sometimes","TRUST, a new value model, flags shaky reasoning in self-driving and robot-arm AI, but fixing the reasoning didn't always fix the robot's actual performance.","Researchers built a system that reads an AI robot's internal reasoning in real time and rewrites the bad parts before it acts.\n\nThe team trained an offline value model called TRUST (Token-level Reward for Utility-Steered Chain-of-Thought) to predict whether a chunk of a vision-language-action policy's chain-of-thought will turn out correct, then uses that score to steer the reasoning as it is generated. Tested on the Alpamayo 1.5 driving policy, TRUST caught correctness lapses with 88.9% accuracy and raised reasoning correctness from 75.9% to 90.0%. In simulated driving scenarios on a challenging subset of AlpaSim, it cut collision rate by 30.4% and trimmed maximum trajectory error by 11.5% versus the unsteered policy, beating a compute-matched Best-of-4 baseline that just samples several outputs and picks the best one. On a separate manipulation policy, DeepThinkVLA, the same approach pushed grasp-state claim accuracy from 69.3% to 90.2% and action-choice accuracy from 68.8% to 85.9% - but closed-loop task performance on the LIBERO-Plus benchmark barely moved.\n\nChain-of-thought is increasingly pitched as a safety feature for embodied AI: if a robot explains itself, the thinking goes, you can catch mistakes before they become crashes. This paper is a useful reality check. Cleaning up the reasoning text measurably helped the driving policy behave more safely, but it did nothing for the robot arm's actual task success, even though its reasoning also got more accurate on paper.\n\nA robot that explains itself more correctly is not automatically a robot that performs better - worth remembering the next time a demo leans hard on its chain-of-thought output.","[\"ai\",\"robotics\",\"self-driving\",\"chain-of-thought\"]","2026-10-02T04:00:00.000Z","2026-10-03T04:33:32.510Z","2026-10-03T04:33:37.770Z","published",null,[],"ai",[24,26,27,28],"robotics","self-driving","chain-of-thought",[30],{"name":31,"url":32},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.00601",0,{"sections":35},[36,39,43,47,52,56,60,65,70,75,80,85,90,95],{"name":37,"slug":24,"count":38,"latest_published_at":18},"AI",5977,{"name":40,"slug":41,"count":42,"latest_published_at":18},"Security","security",842,{"name":44,"slug":45,"count":46,"latest_published_at":18},"Policy","policy",438,{"name":48,"slug":49,"count":50,"latest_published_at":51},"Deals","deals",317,"2026-10-01T22:00:00.000Z",{"name":53,"slug":54,"count":55,"latest_published_at":18},"Hardware","hardware",199,{"name":57,"slug":58,"count":59,"latest_published_at":18},"Science","science",173,{"name":61,"slug":62,"count":63,"latest_published_at":64},"Consumer Tech","consumer-tech",155,"2026-10-01T19:54:10.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Dev Tools","dev-tools",96,"2026-10-01T16:57:03.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Software","software",93,"2026-09-30T21:41:11.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Startups","startups",90,"2026-10-01T21:55:22.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]