AI providers billing by the reasoning path have a financial incentive to pad the count - and a new paper shows how easy that is to hide.
Researchers built an algorithm that generates extra reasoning paths for a self-consistency setup - the technique where a model produces several answers and takes a majority vote - then reorders them so every single path looks necessary to reach that majority. Tested against instruct models from the Llama and Qwen families and reasoning models distilled from DeepSeek-R1 on math, science, and question-answering benchmarks, the algorithm consistently added paths that appeared essential rather than superfluous. The distribution of padded paths turned out to be heavy-tailed, meaning most of the time the padding is modest but occasionally it is severe. Even an auditor built specifically to catch this, tuned to keep its false-positive rate below 10 percent, still let a meaningful amount of overcharging through.
This matters because self-consistency has quietly become a standard way to make LLM answers more reliable, and per-path pricing means more paths equals more revenue for the provider - a conflict of interest nobody was really auditing until now. It is the metered-billing version of a taxi meter running while stuck in traffic: technically defensible, practically unverifiable by the passenger.
Nothing here proves any provider is actually doing this - only that the incentive exists and the audit trail is weaker than it looks.