Security/ llm-security · semantic-caching · cache-poisoning · ai-infrastructure

New Defense Blocks Up to 98% of LLM Cache Poisoning

Researchers found that LLM semantic caches can be tricked by near-duplicate queries, and built a text-based filter that catches most of the fakes.

LLM providers cache answers to save money, and a new paper shows how easily that cache can be poisoned.

Semantic caches work by matching new questions to old ones based on how similar their embeddings look, then serving up the stored answer instead of generating a fresh one. Researchers describe a trick that exploits this: pair a near-duplicate of a legitimate query with extra, malicious content tacked on. The embedding still looks close enough to the real query to trigger a cache hit, but the stored answer reflects the attacker's addition, not the original request. The paper calls this a rewrite-residual pattern, and across three categories of poisoning attacks the researchers tested, it held up consistently.

This matters because caching is now a default cost-saving layer in LLM infrastructure, not an edge case. The paper's framing is useful: embedding similarity measures resemblance, not correctness, and systems that conflate the two inherit a blind spot by design. The proposed fix checks whether deleting part of the cached query's text still produces a stored answer that makes sense, which caught 82 to 98.2 percent of poisoned entries with only a 5 percent false-positive rate and minimal added latency.

It is a reminder that most LLM security problems are old distributed-systems problems wearing new vocabulary. A cache that trusts similarity over provenance is a cache with a trust problem, model embeddings or not.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →