AI/ speech-ai · kv-cache · full-duplex-models · arxiv

New Trick Shrinks Memory Load for AI Voice Chat

Researchers found a way to cut memory use in real-time AI voice models by turning old audio into text instead of deleting it.

AI voice assistants that let you interrupt and talk over them have a memory problem, and researchers just found a fix that does not throw away information to solve it.

Full-duplex speech models, the kind built to listen and talk at the same time, keep a running memory of every sound they hear. That memory, called a KV cache, balloons the longer a conversation runs, which is a real constraint for anything meant to hold a ten-minute call. A team publishing on arXiv built a system that exploits a small window they call listening-time slack, the gap between when the model finishes processing an audio chunk and when the next one arrives. In that gap, a side channel transcribes the speech into text, and once the cache hits a set budget, older raw audio gets evicted while the compact transcript sticks around. Trained with LoRA and distillation to preserve the original model's behavior, the approach cut peak KV-cache size by 64.6 percent in a MiniCPM-o 4.5 implementation on ten-minute sessions.

Memory bloat is the boring but real ceiling on how long an AI voice agent can hold a conversation before it gets slow or falls over, and most fixes so far have meant either capping session length or just deleting old context and hoping nothing important was in it. Swapping raw audio for text is a smarter trade: transcripts are cheap to store and, according to the paper, actually improved transcription, temporal question answering, and summarization rather than just avoiding a hit.

The catch is that this is one lab's benchmark, not a shipped product, and turn-taking behavior on Full-Duplex-Bench came out merely comparable to the uncompressed baseline, not better. Still, if voice interfaces are going to handle real phone-length calls, something like this is the unglamorous plumbing that has to exist first.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →