[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-squeeze-a-975b-parameter-model-onto-11-pcs":10,"sections":45},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":35,"tags":36,"sources":40,"feedback":44,"feedback_at":22,"cost_usd":44,"total_tokens":44},10568,"researchers-squeeze-a-975b-parameter-model-onto-11-pcs","Researchers Squeeze a 975B Parameter Model Onto 11 PCs","A cluster of eleven Intel AI PCs now runs a 975 billion parameter AI model, hitting 60 tokens per second without a data center in sight.","A research team wired eleven ordinary AI PCs together and got a 975-billion-parameter language model to run on them.\n\nThe system, called Cascadia, distributes a mixture-of-experts model named Inkling - 975 billion parameters total, with only 41 billion active per token - across eleven Intel Core Ultra 7 358H machines, each with 64GB of memory, Arc B390 integrated graphics and a gigabit Ethernet connection. A custom engine splits the model's decoder layers six to a machine, compresses them into graphs built for Intel's OpenVINO toolkit, and runs expert computation in FP16 before restoring outputs to FP32. The team says this trimmed the time to execute dense feed-forward blocks from about 8.1 milliseconds to 4.5. Across fifteen concurrency tests, the cluster peaked at 60.29 tokens per second of combined output at 88 simultaneous streams, with median first-token latency of 6.05 seconds at fifteen streams.\n\nMixture-of-experts models like this one are usually the province of data centers stacked with GPUs, not a dozen desktop-class machines on an office network. Cascadia's results suggest the ceiling on serving near-trillion-parameter models is set less by total memory, since each stream could handle context windows up to 512,000 positions, and more by a single-threaded CPU attention bottleneck. That is a solvable engineering problem, not a hardware wall.\n\nCall it a proof of concept, not a product: 60 tokens per second split across 88 users works out to well under one token per second per person, so nobody is replacing their cloud GPU bill with a stack of mini PCs just yet.","[\"ai\",\"mixture-of-experts\",\"edge-computing\",\"research\"]","2026-10-07T04:00:00.000Z","2026-10-08T19:39:12.063Z","2026-10-08T19:39:17.891Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The opening sentence calls the eleven machines 'AI laptops,' but the source and the rest of the draft correctly identify them only as 'AI PCs' — fix that word to match the rest of the piece since the source never confirms laptop form factor.","resolved",{"id":31,"reviewer":32,"round":33,"reason":34,"status":29},"publisher-r2","publisher",2,"The hardware spec is self-contradictory — 'Arc B390 integrated graphics' names a discrete Battlemage-series Arc GPU while the sentence simultaneously claims 'no discrete GPU required,' and the CPU name 'Core Ultra X7 358H' doesn't match Intel's actual Core Ultra 7 358H naming.","ai",[35,37,38,39],"mixture-of-experts","edge-computing","research",[41],{"name":42,"url":43},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.07219",0,{"sections":46},[47,51,56,61,66,71,76,81,86,90,95,100,105,110],{"name":48,"slug":35,"count":49,"latest_published_at":50},"AI",6429,"2026-10-07T18:45:00.000Z",{"name":52,"slug":53,"count":54,"latest_published_at":55},"Security","security",902,"2026-10-07T19:53:42.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Policy","policy",474,"2026-10-07T18:23:21.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":65},"Deals","deals",453,"2026-10-07T23:58:31.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Hardware","hardware",222,"2026-10-07T21:19:54.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Science","science",186,"2026-10-06T21:20:39.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Consumer Tech","consumer-tech",174,"2026-10-07T17:41:41.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Software","software",113,"2026-10-07T18:10:00.000Z",{"name":87,"slug":88,"count":84,"latest_published_at":89},"Startups","startups","2026-10-07T23:36:57.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Dev Tools","dev-tools",105,"2026-10-07T16:59:11.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"General","general",61,"2026-10-07T22:00:24.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"Gaming","gaming",56,"2026-10-07T12:00:00.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"Reviews","reviews",33,"2026-10-05T11:57:17.000Z",{"name":111,"slug":112,"count":113,"latest_published_at":114},"How-To","how-to",8,"2026-10-05T09:00:00.000Z"]