AI/ ai agents · ai safety · benchmarking · research

Researchers Propose a Test for AI Agent Identity Claims

A new framework checks whether an AI agent's claimed memory or model update is backed by real evidence, not just a plausible-looking state.

A new academic framework wants to verify that an AI agent claiming to have updated its memory or switched models is telling the truth.

Researchers introduce the Cognitive Continuity Test (CCT), a checklist-style contract for judging whether a persistent AI agent's claimed change in state - a memory update, a belief revision, a swap to a new underlying model - is actually authorized and real. The test checks five things: who had authority to make the change, where the change came from, whether it was applied in a repeatable way, whether it meets specific logical conditions, and whether there is a tamper-evident receipt linking the old and new states. The team built a benchmark called IdentityLineageBench with 24 families of simulated transitions, then ran a reference verifier against 576 held-out test cases and matched every single label. Two simpler comparison methods, one that just checks how similar the before-and-after text looks and another that checks lineage alone, let through 60 percent and 80 percent of invalid transitions respectively.

That gap is the real finding here: there is currently no standard way to tell whether an AI agent claiming continuity, in effect saying "I'm the same assistant, I just learned something new," actually is, versus being quietly retrained, swapped, or fed a doctored memory. As companies let agents run autonomously for weeks or months, unverified claims about an agent's own state stop being a philosophical question and start being a compliance and trust problem.

The authors are upfront that this is a synthetic benchmark result, not proof the method beats a deployed, policy-aware system, and that detecting actual sleeper-agent behavior or live model migrations is still unmeasured. A 6.21 millisecond median check time across 18,000 runs suggests the approach is cheap enough to try. Whether it holds up outside a generated test set is the next question, not this one.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →