AI/ ai-agents · ai-safety · authorization · self-modifying-ai

A Formal Fix for AI Agents That Rewrite Their Own Code

A new paper proposes rules to stop self-modifying AI agents from duplicating or resurrecting permissions as they fork, replace, or roll back code.

A new arXiv paper tackles a problem most AI safety talk skips: what happens to permissions when an AI agent copies, replaces, or rolls back itself.

The authors propose "authorization succession," a set of rules for tracking authority across generations of self-modifying agents. Each new version gets tied to a manifest, a root, and a full lineage back to its parent, so copies cannot duplicate quotas or inherit access after their ancestor is cut off. The team tested this on 32 registered authorization decisions, each run twice through different interfaces for 64 replays total, matching expected results with 28 allows and 36 denies. A separate independent checker then re-verified all 64 of those original traces and correctly rejected 28 deliberately broken versions, along with passing a batch of crash and scheduling checks. On top of that, two external adapters built for other self-modifying systems, OurArk and the Darwin Godel Machine, reproduced all 32 of the original decisions under real mutation scenarios like process restarts and handoffs.

This matters because self-modifying agents are no longer theoretical. Darwin Godel Machine and similar projects already let software rewrite and redeploy itself, and most permission systems were not built for code that can fork and roll back its own identity. Without something like this, a rolled-back copy could quietly keep access it should have lost.

It is still a proof run on a testbed, not a production system, so treat it as a credential-management fix for a problem the industry has not widely admitted it has yet.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →