A new benchmark finds that AI booking agents don't just search for hotels - they quietly decide how much you pay for them.
Researchers built PriceBench, a diagnostic tool that reverse-engineers an LLM's price, quality, and brand preferences from its hotel picks using a logit choice model. They ran 28 LLMs from 8 providers through 3,600 booking tasks drawn from 179 real New York City hotels. The results split models on consistency, not taste: stronger models held firm, repeatable preferences, while weaker ones either fixated on whichever listing appeared first or picked almost at random. Among models with real preferences, tolerance for price varied by more than an order of magnitude, and the price-versus-quality tradeoff alone moved the average nightly rate from $247 to $393 on identical requests.
That's the part worth sitting with. Hand a booking task to an LLM and you're not getting a neutral search - you're getting whatever purchasing policy that specific model happens to have baked in, unlabeled and undisclosed. Swap one model for another as your travel agent and your booking budget moves with it, for no reason related to the trip itself.
It's the same lesson search engines learned two decades ago about ranking bias, except now the bias sets a dollar figure directly. Releasing the tasks and code is the right instinct - if agents are going to spend our money, the preferences quietly driving that spending shouldn't stay a black box.