Researchers built an audit system that stops AI agents from acting until proof of approval is locked in, and it caught the system's own bug red-handed.
Janus puts the record of an AI agent's actions on the same path as the action itself. Before a step can run or its effect be released, the proposal, the verdict, and any human or validator sign-off must be durably logged in a signed, hash-chained record, so an auditor can later re-derive every verdict offline using just the log and a public key. The team tested it under crash injection (144 in-process kills, 81 through a daemon) and verified a 100-million-event log offline in about 254 seconds. They then ran a real model behind a simulated lending workflow, comparing a Janus-governed agent against an ungoverned one on identical recorded outputs: when the lending mandate wasn't baked into the model's own prompt, the model approved six loans over the limit, the plain agent paid out all six, and Janus paid out none, each one caught by a deterministic validator.
The real story is what the testing surfaced. Building the experiment exposed a bug in Janus's own approval logic: a human sign-off meant for one loan request got silently counted for a different proposal, inflating an approved amount from 100 to 1,000,000. The researchers disclosed it themselves, along with a first fix and five mitigations, and were explicit about what the system still doesn't guarantee.
As AI agents get handed real financial authority, the gap between logging an action and proving it after the fact is exactly where accountability breaks down, and this team found that gap in their own supposedly audited system before anyone else did.