Researchers built a credit-scoring AI agent that can patch its own rulebook when regulations shift, but only through a monitored process that writes every change to an audit trail before it ships.
The system, described in a new arXiv paper, draws a hard line: the AI's underlying model weights never change. Only the surrounding harness (its instructions, tool-calling logic, and how it composes basic operations) can be rewritten, and every edit must pass an admission gate that logs a hash-chained record before deployment. In simulation, using a scripted agent and a seeded-search proposer rather than actual language models, the gated version approved just 144 of 7,449 proposed changes across three types of regulatory re-interpretation, and none of the accepted changes made error rates worse on historical data. A looser gate that mimicked how unsupervised systems typically self-check, by just watching for fewer errors in recent cases, let through 309 harmful changes and left the miss rate for risky loans above 10% in 49 of 90 test runs.
The findings map directly onto the EU AI Act's high-risk credit-scoring provisions, which already demand documented, auditable changes to lending algorithms. They also expose a harder limit: the gate correctly rejected every proposed fix when asked to relabel historical cases under a new rule, and it only half-solved the toughest structural shifts, blocking a valid fix in half the test seeds. That is a real ceiling on how much self-correction a compliance-bound system can manage before a human has to step in.
The US is not ignoring this fight; its April 2026 model-risk guidance already covers algorithmic lending, but that guidance explicitly excludes agentic AI, leaving the exact systems this paper is trying to tame outside its reach.