AI/ ai agents · ai safety · autonomous systems · research

Researchers Propose Rules for When AI Agents Lose Authority

A new paper argues AI agent trust should be a binary contract, not a score, and tests the idea with failure-injection experiments in coding tasks.

A new academic framework wants AI agents to lose authority automatically the moment the evidence turns bad, instead of running on a trust score that can be argued away.

Researchers published a paper proposing a 'Runtime Assurance Contract,' or RAC, a formal policy schema that ties an AI agent's permissions to the quality of available evidence. Under RAC, a failed or unknown mandatory check forces a retry, an escalation to a human, or a full stop, no matter how good the agent's aggregate performance score looks. The team tested the idea with a failure-injection study of 280 constructed cases in agentic coding, comparing a gate-based system against a score-only rule and a restricted baseline protocol. At the paper's published weights, the score-only rule caught 80 of 100 cases that should have been blocked and all 40 that needed review; tuned after the fact, it matched the gate system exactly. A separate check used 18 hand-authored traces to test whether version-pinned evidence and review transitions held up against simpler policy variants. In a further prospective test of 24 synthetic episodes, two blinded LLM judges agreed on how to label all 72 action attempts made during that holdout, and both RAC and a separately built stateful baseline matched those judgments.

The paper's real point is that 'trustworthy enough on average' is a bad standard for agents doing consequential work; a single unresolved red flag should outweigh a good aggregate score. That's a useful corrective as companies hand AI agents more control over code, money, and decisions without a clear mechanism for cutting them off mid-task.

Worth noting: this is still a synthetic-data exercise with no production deployment behind it, so the real test is whether anyone building commercial agents wires in an actual hard stop instead of just logging a warning and moving on.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →