AI/ ai · prompt-optimization · llm-agents · benchmarks

Adaptive-GEPA Teaches AI Agents to Route Their Own Tasks

A research team split one AI workflow into a router plus specialist sub-programs, beating two rival optimization methods on a four-task benchmark.

A new optimizer teaches AI systems to split incoming requests into specialties instead of forcing one script to handle everything.

Researchers built Adaptive-GEPA, an extension of the prompt-optimization framework GEPA, which rewrites prompts and code based on execution traces and evaluator feedback. Instead of tuning a single shared program, Adaptive-GEPA evolves a router alongside a library of specialist programs, all under one shared search budget. The router's instructions, each specialist's description, and its code are edited as plain text based on feedback, and specialists that handle overlapping requests get merged by combining their descriptions and code. In a test with Qwen3-8B across four task families, the system grew four specialists on its own, without being told what the task categories were, and its routing matched the correct task split on all 651 test requests.

The headline number: family-mean test scores rose from 52.6 to 70.6 out of 100, comfortably ahead of GEPA's single full-program approach (62.5) and GRPO, a reinforcement-learning fine-tuning method, which scored 54.0 at a nominal budget of 18,000 scored calls. That gap matters because most teams running an AI agent behind one API endpoint get a flood of different request types, and the usual fix is manually building separate pipelines for each. This suggests letting the system discover and divide that work itself can beat both a one-size-fits-all prompt and a heavier retraining approach.

Worth noting: the paper's own caveat is that these call counts don't equate to total compute, so the efficiency claim deserves a second look before anyone swaps out their routing logic.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →