[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-propose-rules-for-when-ai-agents-lose-authority":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},8956,"researchers-propose-rules-for-when-ai-agents-lose-authority","Researchers Propose Rules for When AI Agents Lose Authority","A new paper argues AI agent trust should be a binary contract, not a score, and tests the idea with failure-injection experiments in coding tasks.","A new academic framework wants AI agents to lose authority automatically the moment the evidence turns bad, instead of running on a trust score that can be argued away.\n\nResearchers published a paper proposing a 'Runtime Assurance Contract,' or RAC, a formal policy schema that ties an AI agent's permissions to the quality of available evidence. Under RAC, a failed or unknown mandatory check forces a retry, an escalation to a human, or a full stop, no matter how good the agent's aggregate performance score looks. The team tested the idea with a failure-injection study of 280 constructed cases in agentic coding, comparing a gate-based system against a score-only rule and a restricted baseline protocol. At the paper's published weights, the score-only rule caught 80 of 100 cases that should have been blocked and all 40 that needed review; tuned after the fact, it matched the gate system exactly. A separate check used 18 hand-authored traces to test whether version-pinned evidence and review transitions held up against simpler policy variants. In a further prospective test of 24 synthetic episodes, two blinded LLM judges agreed on how to label all 72 action attempts made during that holdout, and both RAC and a separately built stateful baseline matched those judgments.\n\nThe paper's real point is that 'trustworthy enough on average' is a bad standard for agents doing consequential work; a single unresolved red flag should outweigh a good aggregate score. That's a useful corrective as companies hand AI agents more control over code, money, and decisions without a clear mechanism for cutting them off mid-task.\n\nWorth noting: this is still a synthetic-data exercise with no production deployment behind it, so the real test is whether anyone building commercial agents wires in an actual hard stop instead of just logging a warning and moving on.","[\"ai agents\",\"ai safety\",\"autonomous systems\",\"research\"]","2026-10-01T04:00:00.000Z","2026-10-01T12:44:41.241Z","2026-10-01T12:44:47.366Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The body sentence conflates two distinct validation studies — split it to clarify that the 18 hand-authored traces checked version-pinned evidence\u002Freview transitions against simpler policy variants, while the 'agreed on all 72 judged action attempts' result applies only to the separate 24-episode synthetic holdout, since as written it implies both studies produced that 72-attempt agreement figure.","resolved","ai",[32,33,34,35],"ai agents","ai safety","autonomous systems","research",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.39717",0,{"sections":42},[43,47,52,57,62,67,72,77,82,87,92,97,102,107],{"name":44,"slug":30,"count":45,"latest_published_at":46},"AI",5739,"2026-10-02T04:00:00.000Z",{"name":48,"slug":49,"count":50,"latest_published_at":51},"Security","security",829,"2026-10-01T17:31:56.000Z",{"name":53,"slug":54,"count":55,"latest_published_at":56},"Policy","policy",437,"2026-10-01T18:10:00.000Z",{"name":58,"slug":59,"count":60,"latest_published_at":61},"Deals","deals",317,"2026-10-01T22:00:00.000Z",{"name":63,"slug":64,"count":65,"latest_published_at":66},"Hardware","hardware",198,"2026-10-01T17:38:48.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":71},"Science","science",168,"2026-10-01T18:35:55.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Consumer Tech","consumer-tech",155,"2026-10-01T19:54:10.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":81},"Dev Tools","dev-tools",96,"2026-10-01T16:57:03.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Software","software",93,"2026-09-30T21:41:11.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Startups","startups",90,"2026-10-01T21:55:22.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":108,"slug":109,"count":110,"latest_published_at":111},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]