AI/ ai · llm · multi-agent-systems · research

New Method Lets Different AI Models Share Cache, Skip Re-Reading

A new technique called HeteroFold lets AI models from different families share cached context, cutting redundant reprocessing in multi-agent systems.

A new technique lets AI models from different vendors trade pre-processed context directly, skipping the step where each one rereads everything from scratch.

Researchers describe HeteroFold, a method for transferring a model's key-value cache, the internal representation of everything it has already read, to a completely different model family without retraining either one. The method aligns the two models' internal structures, maps the sender's cache into a format the receiver understands, and calibrates it so the receiver's behavior does not degrade. Tested across six model-to-model transfer directions, it beat prior cache-transfer methods on four long-context benchmarks and most short-context tests, and matched plain text communication on a multi-agent benchmark. In one test case, moving context from Llama-3.1-8B to Ministral-3-14B at 32K tokens ran 10.7 times faster than standard prefill, and 1.18 to 1.47 times faster than the best existing prefill-free alternatives.

The real target here is multi-agent AI systems, which increasingly mix models from different makers for different jobs, one for planning, another for coding, another for retrieval. Today, handing a task from one model to another usually means the receiving model re-reads all the shared context from the start, burning time and compute on work already done. Removing that redundancy without touching either model's weights is what makes this useful outside a research lab.

It is still a lab result tested on specific model pairs, not a drop-in standard, so treat the speedup numbers as a ceiling rather than a guarantee for your own stack.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →