AI chatbots are bad at remembering why they made a decision earlier in a long chat, and a new benchmark proves it with numbers.
Researchers built SCALE-QA, a 3,000-question test that drops ordinary task requests into long, mixed-topic conversation threads where there is no helpful "new session" marker. To answer correctly, a model has to find and use a specific piece of evidence from earlier in the thread, not just recall a fact about the user. The team ran the test through 128k-token contexts, plus a 400-question version at 1M tokens, and tested it against both retrieval-augmented systems and long-context models. They also built their own fix: TSIM, which chops the conversation into distinct episodes and indexes them in a multi-layered memory system instead of just stuffing more tokens into the context window.
This matters because most memory benchmarks test a strawman: a chatbot that already knows which chunk of conversation holds the answer. Real usage is messier. People bounce between topics in one thread, and a model has to figure out, unprompted, which earlier exchange actually justifies its next answer. SCALE-QA targets that gap directly, and the result is that long-context window size alone does not solve memory - structure does.
TSIM beat the best baseline by 5.6 to 17.6 accuracy points across three different LLM backends, which is a meaningful gap for anyone building assistants meant to hold a conversation for more than a few turns. It is also a quiet rebuttal to the industry's favorite flex of the past two years: bigger context windows. A 1M-token window is not memory if the model cannot tell which million tokens matter.