[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-ai-memory-system-nearly-matches-rival-audit-finds-cracks":10,"sections":44},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":34,"tags":35,"sources":39,"feedback":43,"feedback_at":22,"cost_usd":43,"total_tokens":43},8662,"new-ai-memory-system-nearly-matches-rival-audit-finds-cracks","New AI Memory System Nearly Matches Rival, Audit Finds Cracks","A new long-term memory system for AI models scores near a published benchmark leader, but its own grading judge contradicts itself on identical answers.","Researchers have published a fully auditable long-term memory system for AI chatbots, and it nearly matches the best published score on a standard memory test.\n\nThe system, tested on the LongMemEval-S benchmark (a 500-question test of how well an AI remembers details across long conversation histories), skips black-box memory tricks for a documented retrieval chain: it pulls candidate information with hybrid search, reranks it with a cross-encoder, assembles an evidence packet designed to cover the full answer, and only then hands that packet to a large language model to write the final response. Using Claude Opus as that final reader, two separate 500-question runs scored 479 and 475 correct under GPT-4o grading, straddling the 478 score published for a rival system called Chronos High. Because the two systems differ in reader model version, grading prompt, and possibly the underlying data version, the researchers say the results don't prove their system is better or worse, just comparable. Swapping in other reader models moved the needle in both directions: xAI's grok-4.6-high scored 476 and 474, while a version set to maximum reasoning effort actually did worse, dropping to 461 and 465.\n\nThe real story is in the fine print. The same GPT-4o judge flipped three of its own verdicts when asked to rescore byte-identical answers from the first run, a reminder that even the yardstick used to measure these systems isn't fully consistent. A second, independent judge mostly agreed with the primary one, matching on 493 of 500 rows, but still landed on a lower score of 472 for both passes, underscoring how much of these benchmark wins depend on which judge, and which day, you ask.\n\nAll of this was built and tuned on the same 500 questions it was tested on, with no held-out set and no human double-checking the judge, so treat the near-parity with Chronos High as a data point, not a verdict.","[\"ai\",\"benchmarking\",\"long-term-memory\",\"llm-evaluation\"]","2026-09-30T04:00:00.000Z","2026-09-30T17:55:58.809Z","2026-09-30T17:56:01.822Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Name and date the source explicitly (e.g. the arXiv preprint arXiv:2609.38021) instead of the vague 'a new paper' — as written there's no named publication, venue, or posting date backing the cited figures.","resolved",{"id":31,"reviewer":26,"round":32,"reason":33,"status":29},"editor-r2",2,"Fix the misattribution in paragraph three: per the source, it's the primary\u002Fofficial GPT-4o judge that flips three verdicts on byte-identical pass-1 answers, not the second judge — as written the draft wrongly implies the second judge contradicts itself.","ai",[34,36,37,38],"benchmarking","long-term-memory","llm-evaluation",[40],{"name":41,"url":42},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.38021",0,{"sections":45},[46,49,53,57,62,67,71,76,81,85,90,95,100,105],{"name":47,"slug":34,"count":48,"latest_published_at":18},"AI",5180,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Security","security",791,{"name":54,"slug":55,"count":56,"latest_published_at":18},"Policy","policy",417,{"name":58,"slug":59,"count":60,"latest_published_at":61},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":63,"slug":64,"count":65,"latest_published_at":66},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":18},"Science","science",155,{"name":72,"slug":73,"count":74,"latest_published_at":75},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":86,"slug":87,"count":88,"latest_published_at":89},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]