Researchers just tested how ChatGPT, Claude, Grok, and DeepSeek decide to search the web, and it turns out there's no shared playbook.
The study, the first to trace the full lifecycle of agentic web search across four major conversational platforms, combined real user interaction logs with controlled experiments run through each platform's own API. The researchers tracked when each agent chose to search, how it phrased its queries, which domains its results favored, and how it turned those results into a final answer. They found that the decision to search at all varies widely by platform and model, and that searching more often doesn't translate into better responses. Each platform also uses its own distinct query strategy, and each one's search engine tends to surface results skewed toward a particular set of preferred domains.
That matters because these tools are increasingly treated as a substitute for a search engine, complete with an implied promise that answers are backed by real, checkable sources. The study found responses are mostly grounded in what the search actually returned, but some claims cited no source at all, which is exactly the kind of gap that turns a helpful assistant into a confident-sounding guesser. Add in the domain preferences baked into each platform's search engine, and the "neutral" answer you get depends more on which chatbot you asked than you might assume.
It is the AI equivalent of learning that four different reporters, sent to cover the same press conference, would each ask different questions, trust different sources, and write different stories, with only some of them bothering to say where their quotes came from.