Apple's rebuilt Siri runs on a 1.2-trillion-parameter model built on Google's Gemini technology, hosted on Google Cloud, and processed on Nvidia Blackwell B200 GPUs.
At WWDC 2026, Apple confirmed the architecture rather than obscuring it. The new Siri backend is not a proprietary model trained from scratch — it relies on Gemini foundations and runs on infrastructure owned by the company Apple competes with directly in smartphones. Nvidia's Blackwell B200 GPUs handle the compute load. Apple devoted notable keynote time to arguing that none of this arrangement compromises user privacy.
The problem is that Apple's privacy positioning has long functioned as a quiet rebuke of Google's data-collection model — the implication being that what happens on your device stays there. Running Siri inference on Google Cloud makes that contrast harder to sustain without explanation. The architecture also reveals something Apple rarely concedes: that it could not field a competitive trillion-parameter model entirely on its own infrastructure.
A company that feels genuinely comfortable with its privacy architecture does not typically spend keynote time explaining it.
