AI/ llm-serving · kvcache · moonshot-ai · infrastructure

Moonshot's Mooncake Serves Kimi With Idle CPU and SSD

Moonshot's Mooncake splits LLM serving and reuses idle GPU-server storage, lifting Kimi's capacity by 75 percent.

Moonshot AI rebuilt how its Kimi chatbot serves answers, and the fix was in storage, not more chips.

The company's new system, called Mooncake, splits the two stages of answering a prompt, reading the question (prefill) and generating the reply (decoding), onto separate clusters instead of running them together. It then taps spare CPU, memory, and SSD capacity already sitting idle on GPU servers to store KVCache, the running record of previously processed tokens that would otherwise need to be recomputed from scratch. A KVCache-aware scheduler decides where to route each request, trying to maximize throughput while keeping response times inside set targets. When traffic spikes past what the system can handle, Mooncake predicts which requests it will fail to serve in time and rejects them early instead of letting them clog the queue.

Long-context requests, the kind that choke most LLM serving stacks, are where the gains show up most: Mooncake reported up to a 525% throughput increase over its baseline in simulated tests while still hitting latency targets. In production, Moonshot says the architecture let Kimi handle 75% more real requests on the same hardware, a result from reorganizing infrastructure rather than swapping in a bigger model.

Most serving stacks treat GPU memory as the only resource worth squeezing. Mooncake's bet, that idle CPU, DRAM, and SSD sitting next to the GPUs are worth harvesting too, is a cheaper lever than buying more accelerators, and a sign of how tight GPU supply has made infrastructure efficiency its own competitive edge.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →