AI/ ai · llm inference · cloud providers · research

Picking an AI Model Is Only Half the Battle, Study Finds

A new arXiv study finds AI provider price does not predict quality, speed, or uptime, and proposes a router that avoids bad picks.

Open-weight AI models are commodities now, but the companies serving them are not created equal, according to new research.

In a paper titled 'You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference' (arXiv:2609.37902), researchers measured live API endpoints across multiple open-weight models, competing providers, and task types over three separate measurement waves. They found that the same model varies sharply in quality, latency, and uptime depending on which provider serves it. Pricier providers were consistently faster, but price alone did not predict quality or availability. One provider's deployment handled knowledge questions almost perfectly but fell apart on multi-step reasoning tasks.

Most routing tools today only decide which model to call, treating every provider hosting that model as interchangeable. This paper argues that is the wrong assumption: provider choice is its own decision, and getting it wrong can mean paying a premium for a broken endpoint or saving money by accepting silent quality loss. The authors propose FACET, an online router that certifies which provider-task combinations are safe before routing to them, and defaults back to a trusted anchor provider when it is not sure.

The paper has not been peer-reviewed, but if the pattern holds, the open-weight model market is less like a single product with many resellers and more like a used-car lot where the sticker price tells you almost nothing about what you are actually getting.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →