A research team audited their own AI companion and found its memory system was mostly broken, even though users rated it as working fine.
The researchers built and deployed a proactive AI companion called Lita for a month with nine colleagues, then set its self-descriptions against user judgments and implementation records, grading each claim as supported, contradicted, or unresolved. Participants agreed with Lita's claims about its personality and style, but did not endorse its claims about the relationship it said it had built with them, and rated its memory at or above the midpoint. Checking the code, the researchers found two of the agent's three memory layers had never actually executed the step that would store new information.
That gap matters beyond one small pilot. Companion apps increasingly sell themselves on remembering who you are and how you've changed together, and users have no way to check those claims from the outside. This audit shows that fluent self-description and decent user ratings are not proof the underlying mechanism ran.
The researchers built this system themselves and still needed a formal audit to catch the problem, so companies selling companion AI without that kind of internal check should not get the benefit of the doubt.