AI/ prompt-engineering · llm-optimization · ai-research · automatic-prompt-optimization

Researchers Build Prompt Optimizer That Works on a Budget

BudgetAPO optimizes prompts with a fraction of the API calls rival methods need, and rarely gives up and returns the original prompt unchanged.

A new prompt-optimization method called BudgetAPO is built for a world where every API call costs money.

Researchers introduce BudgetAPO, a single-stage optimizer designed for tight-budget settings, where only a small number of calls to the subject model are available. It sizes each evaluation slice to a task's noise level using a short probe, locks in a fixed slice so every accept-or-reject decision becomes a paired comparison, and uses a reflective operator that rewrites reasoning strategy and output format together rather than separately. The team tested it across seven benchmarks and five subject models against established baselines like GEPA and OPRO.

The gap that matters here is practical, not academic. GEPA and OPRO assume hundreds to thousands of subject-model calls, which is a reasonable lab assumption and a terrible one behind paid, rate-limited APIs. At a 250-call budget, GEPA returned the original seed prompt unchanged 86% of the time - meaning it burned the budget and accomplished nothing. BudgetAPO did that in only 13% of runs, and on GPT-OSS-20B, GEPA needed five times as many calls to match what BudgetAPO achieved at just 100.

Most prompt-optimization papers are benchmarked like compute is free. This one is a useful corrective: for teams actually paying per call, an optimizer that knows when to stop probing and when to commit is worth more than one that scores marginally higher given unlimited tries.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →