AI/ ai · llm-memory · gpu-efficiency · research

New AI Memory System Skips Recomputing 50 Million Tokens

A new disk-based memory layer lets AI models recall facts from 50 million tokens back without recomputing them, cutting GPU energy use by up to 12 times.

A new memory layer for AI models can recall facts from 50 million tokens back without ever recomputing them.

Researchers tested a tool called galahad-kv that saves a model's internal key-value state to encrypted local NVMe disk in chunks of about 16,000 tokens, then reloads those chunks later instead of reprocessing the text. Running on a single Nvidia H100 GPU with Gemma 3 models at 12B and 27B parameters, the system pulled back every one of 100 tested memory blocks correctly, byte for byte, across a 50-million-token stream. Loading a stored block took roughly a third to a quarter of the time of recomputing it, and it used 8.8 to 12.3 times less GPU energy. GPU memory usage stayed flat the entire time, no matter how deep into the 50-million-token history the system reached.

Context windows have grown mostly through brute-force scaling, but every token in them still costs compute to reprocess on each query. This approach treats memory as storage rather than something to recalculate, which is why it is cheaper and faster, though it only works for text the model has already seen once. The larger model, Gemma 3 27B, answered correctly 98 times out of 100 when quizzed about facts buried deep in that history; the smaller Gemma 3 12B got 82 right, and neither model made up an answer.

The catch is the storage bill: keeping 50 million tokens of state on standby takes terabytes of local disk, so this trades one expensive problem, compute, for another, disk space - not a free lunch, just a different one.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →