A new research paper hands AI inference clusters a way to skip the dedicated traffic cop.
Cascadia is a system for running large language models across fleets of commodity Intel AI PCs, tapping each machine's CPU, integrated GPU, and NPU. Instead of routing every request through a central scheduler, each node handles its own ingress, scheduling, and execution. Nodes find each other over a libp2p QUIC mesh, prove they belong using CA-issued ed25519 certificates, and gossip live load data so requests land on whichever peer can handle them. The system can run a model whole on one node, spread it across load-balanced replicas, or split it into a pipeline-sharded chain, and it logs signed, hash-chained receipts for every response so results can be audited later.
The researchers tested a three-node Phi-3.5-mini setup on NPUs and measured 3.1 times the response throughput of a single node under ten concurrent requests; a separate four-node deployment hit 4.06 times single-node throughput. That matters because the alternative today is hyperconverged infrastructure from vendors like IBM, Nutanix, VMware, and HPE, which typically demands dedicated control-plane hardware, specialized licensing, and purpose-built racks just to coordinate AI workloads.
Three and four nodes is a lab bench, not a data center, and the vendor comparison in the paper leans on published documentation rather than head-to-head testing, so treat the speedup numbers as a promising first measurement, not a verdict.