AI/ ai · llm-inference · caching · research

Galahad Memory Layer Makes LLMs Stop Rereading Documents

A new caching layer lets language models skip rereading documents they've already seen, cutting response times and energy use by roughly 90 percent in tests.

A new memory layer stops language models from rereading documents they have already seen, turning a per-question cost into a one-time one.

Researchers built Galahad, a memory layer for vLLM, SGLang and llama.cpp that pairs two components: Taliesin, which stores a block of text's key-value attention state after the first read and reloads it bit-for-bit instead of recomputing it, and Blaise, which keeps full documents in storage and hands the model only the slice a question needs. In a test with 100 facts buried in a 97,000-token document, a 31-billion-parameter Gemma model running on llama.cpp answered 98 of 100 questions correctly using Taliesin alone, in 3.0 seconds and 572 joules per question, versus 10 of 100 in 9.3 seconds and 2,754 joules without it, because unmodified serving could hold only the last 12,000 tokens of context. Adding Blaise pushed accuracy to 100 of 100 across all three serving engines, at 0.59-0.64 seconds and 200-213 joules per question - roughly a tenth of the baseline's time and energy - while a tuned RAGFlow retrieval pipeline managed only 77.

Most LLM APIs today treat every request as a blank slate, recomputing a document's internal state each time someone asks a follow-up question - a waste the researchers measured at 98.7% of prompt tokens across seven real-world datasets. If document-level caching like this holds up outside the lab, it reframes the economics of serving: reading a document becomes a one-time cost, paid off after about 13 questions by the paper's own accounting, rather than a tax on every question asked about it.

vLLM and SGLang already cache repeated prompt prefixes; what is new here is treating whole documents as reusable, bit-identical state rather than just a shared prefix - and the claim that the system fails closed, recomputing rather than serving any cached load that does not check out, is the part worth watching once this meets real traffic rather than a lab benchmark.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →