AI agents still can't reliably turn their skills into cheap, reusable tools - and a new benchmark proves it.
Researchers built BOTTLED, a benchmark that hands AI agents an entire unlabeled workload and forces them to find their own solution under fixed time, compute and API budgets. The agents could train a small model, write a reusable program, or pick any approach they wanted. Tested across ten models and three tasks, the results were rough: 48 of 60 runs scored below the confidence interval of the model's own zero-shot performance, and 31 of 60 runs lost to a basic small-model distillation baseline using the same token budget. Being good at a task zero-shot, it turns out, does not predict being good at packaging that skill cheaply.
Not every result was bad. On a query-product relevance classification task, Opus 5 kept about 82% of its zero-shot accuracy score (macro-F1) while cutting reported costs roughly 657-fold. It also matched 94% of the performance of Jev, a model built specifically for cheap, repetitive inference, while running at a quarter of Jev's projected cost.
That gap matters because most real-world LLM use is not one clever query - it is millions of near-identical ones, where API costs compound fast. An agent that can reliably distill its own skill into something cheaper would change the economics of deploying LLMs at scale.
But the headline number here is the failure rate, not the win: four out of five bottling attempts underwhelmed, and over half lost to a dumb baseline. Agents shrinking themselves into cheap, scalable tools is still more promise than product.