AI/ perplexity · ai · inference · cloud

Perplexity Routes AI Queries Between Your PC and the Cloud

The search startup's new hybrid inference system offloads simpler queries to local hardware in real time, cutting cloud costs at scale.

Perplexity AI has built a routing layer that sends each AI query to either your PC or a data center, depending on what the task actually requires.

CEO Aravind Srinivas announced the system at Computex in Taipei, calling it an "air-traffic controller" for compute. The platform monitors incoming queries in real time and decides whether a local processor can handle the work or whether it needs to escalate to cloud hardware. Lighter requests stay on-device; heavier inference gets routed to the data center. The stated goal is to handle more query volume without scaling infrastructure costs at the same rate.

The economics are the real story. Cloud inference is not cheap at scale, and offloading even a fraction of routine queries to user hardware chips away at Perplexity's per-query costs in a way that benefits the company at least as much as the user. The architecture also puts Perplexity in the same conceptual territory as Apple Intelligence, which uses a similar on-device-first, private-cloud-second model. The difference is that Apple controls the hardware, the OS, and the cloud endpoint. Perplexity controls only the software layer sitting on top of other people's silicon.

The "air-traffic controller" framing is tidy marketing. What the announcement left open is how the system behaves when local hardware varies widely across Perplexity's user base, and whether users will be told when a given query leaves their device.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →