A new benchmark shows large language models can write optimization code that gets the right answer and still runs painfully slow.
Researchers built OptTips, a knowledge base of 50 expert modeling techniques across eight families, then used a multi-agent framework called OptDachshund to turn existing optimization benchmarks into paired tasks: one solved the standard way, one solved with expert techniques, both scored for correctness and solve time. The result is EfficientOpt, a set of 561 expert-reviewed tasks with reference solutions. Across 11 LLMs tested, the pattern held: when a model's code produced the correct objective value, it typically took longer to solve than the expert-built version of the same problem. Even narrowing to cases where the LLM's model used fewer variables and constraints than the reference, 57% still solved slower.
This matters because most LLM-for-math evaluations stop at did it get the right number. This benchmark shows that passing that bar can still leave you with code that is technically correct and practically expensive - slow to prepare data, slow to build the model, slow to solve, or some combination of the three. For anyone wiring LLMs into scheduling, routing, or resource-allocation systems where solve time is a real cost, correctness alone is not the metric that matters.
Getting the right answer eventually is a low bar. Getting it before your cloud bill notices is the one that counts.