A research team wired eleven ordinary AI PCs together and got a 975-billion-parameter language model to run on them.
The system, called Cascadia, distributes a mixture-of-experts model named Inkling - 975 billion parameters total, with only 41 billion active per token - across eleven Intel Core Ultra 7 358H machines, each with 64GB of memory, Arc B390 integrated graphics and a gigabit Ethernet connection. A custom engine splits the model's decoder layers six to a machine, compresses them into graphs built for Intel's OpenVINO toolkit, and runs expert computation in FP16 before restoring outputs to FP32. The team says this trimmed the time to execute dense feed-forward blocks from about 8.1 milliseconds to 4.5. Across fifteen concurrency tests, the cluster peaked at 60.29 tokens per second of combined output at 88 simultaneous streams, with median first-token latency of 6.05 seconds at fifteen streams.
Mixture-of-experts models like this one are usually the province of data centers stacked with GPUs, not a dozen desktop-class machines on an office network. Cascadia's results suggest the ceiling on serving near-trillion-parameter models is set less by total memory, since each stream could handle context windows up to 512,000 positions, and more by a single-threaded CPU attention bottleneck. That is a solvable engineering problem, not a hardware wall.
Call it a proof of concept, not a product: 60 tokens per second split across 88 users works out to well under one token per second per person, so nobody is replacing their cloud GPU bill with a stack of mini PCs just yet.