AI/ ai · benchmarks · llm-memory · conversational-ai

New Benchmark Finds AI Chatbots Lose Track of Long Conversations

A new test called SCALE-QA shows chatbots struggle to connect task decisions to evidence buried earlier in sprawling, topic-jumping conversations.

AI chatbots are bad at remembering why they made a decision earlier in a long chat, and a new benchmark proves it with numbers.

Researchers built SCALE-QA, a 3,000-question test that drops ordinary task requests into long, mixed-topic conversation threads where there is no helpful "new session" marker. To answer correctly, a model has to find and use a specific piece of evidence from earlier in the thread, not just recall a fact about the user. The team ran the test through 128k-token contexts, plus a 400-question version at 1M tokens, and tested it against both retrieval-augmented systems and long-context models. They also built their own fix: TSIM, which chops the conversation into distinct episodes and indexes them in a multi-layered memory system instead of just stuffing more tokens into the context window.

This matters because most memory benchmarks test a strawman: a chatbot that already knows which chunk of conversation holds the answer. Real usage is messier. People bounce between topics in one thread, and a model has to figure out, unprompted, which earlier exchange actually justifies its next answer. SCALE-QA targets that gap directly, and the result is that long-context window size alone does not solve memory - structure does.

TSIM beat the best baseline by 5.6 to 17.6 accuracy points across three different LLM backends, which is a meaningful gap for anyone building assistants meant to hold a conversation for more than a few turns. It is also a quiet rebuttal to the industry's favorite flex of the past two years: bigger context windows. A 1M-token window is not memory if the model cannot tell which million tokens matter.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →