A new attack shows you can cut in line for an AI model's response just by lying to the scheduler.
Researchers built an attack called JIL that targets prediction-based schedulers used in large language model serving, the kind that prioritize requests expected to need fewer output tokens so the whole system runs faster. JIL works by appending an adversarial suffix to a prompt, tricking a lightweight length-prediction model (tested against a system called TRAIL) into drastically underestimating how long the response will actually be. Across two datasets and four different LLMs, the attack cut predicted output length by as much as 83.4 percent, and manipulated requests finished up to 1.53 times faster on average than they should have. Notably, the drop in predicted length was far bigger than the drop in actual output length, meaning the model kept generating roughly the same amount of text, it just lied about it ahead of time.
This matters because shared AI infrastructure increasingly leans on shortest-job-first style scheduling to keep response times low for everyone, and that only works if requests report their expected length honestly. JIL shows that assumption doesn't hold under adversarial pressure, turning a scheduling optimization into a way for one user to hog compute at everyone else's expense, with the trade-off of inconsistent response quality depending on the model and task.
The researchers also tested a defense: grouping predicted lengths into coarse buckets instead of trusting exact numbers, which blunted the attack without punishing honest requests, a reminder that precision and resilience don't always come as a package deal.