Researchers have built a system that decides, mid-task, which steps of an AI workflow are safe to skip.
The technique, called Learning What to Skip (LW2S), targets multi-agent large language model workflows that chain together planning, execution, verification, and summarization steps. Running every component on every task wastes computation and can even let a later step overwrite an already-correct answer. LW2S treats the decision to skip a step as a credit-assignment problem: it compares full-workflow logs against controlled experiments where a step was deliberately omitted, then learns a safety model for each type of skip. If an early skip looks risky, the controller does not just guess: it keeps running and can reconsider a later step instead.
Tested on math reasoning, multiple-choice QA, and code generation across two instruction-tuned model families, LW2S cut token costs while matching or beating the accuracy of running the full workflow every time. That matters because multi-agent pipelines are becoming the default way to squeeze more reliability out of LLMs, and every unnecessary planning or verification pass is pure cost with no guaranteed benefit.
The paper's own experiments hint at the catch: models agreeing with each other is not enough to justify a skip, which suggests plenty of today's multi-agent setups are paying for verification steps that verify nothing.