A team of researchers has designed an auction system that lets AI users bid for faster responses from language models, instead of being stuck in one of a handful of fixed-price tiers.
The paper proposes an inference auction that allocates priority access to large language model APIs based on how much each request's user is willing to pay for speed. The mechanism is built to incentivize truthful bidding, meaning users gain nothing by lying about how urgently they need a response. The researchers also built an autobidding agent that lets a user set a total budget and let the system automatically adjust bids over time to get the best service within that spending limit. In experiments, the auction increased overall system welfare while preserving the caching efficiency and low latency of SGLang, a widely used open-source inference serving framework.
Right now, API providers mostly sell speed in blunt packages - a standard tier, a priority tier - that force very different users into the same bucket. An auction lets an algorithm that needs an answer in milliseconds pay more than someone running an overnight batch job, in theory squeezing more value out of the same GPUs without buying more of them.
It is still a research design, not something live cloud APIs use today, and anyone who has watched spot-instance prices spike during an AWS outage knows that letting the market set the price can get ugly fast for the user with the smallest budget.