A new framework strips AI research agents of memory between steps - and gets better results for a fraction of the tokens.
Researchers built Stateless Language Agents (SLAs), where no agent carries its own conversation history. Instead, a harness owns all research state - the candidate solutions and measured results - and builds a fresh, role-specific context for every invocation. A stateless 'Advisor' reviews harness-summarized evidence across search directions and assigns specific experiments to parallel 'Worker' agents. Tested against three existing agent frameworks on software engineering, kernel optimization, and algorithm design, at budgets up to one billion tokens, SLA beat all three on every task and matched the strongest kernel-optimization baseline using more than 84 percent fewer tokens, while the Advisor itself consumed less than 0.6 percent of total tokens.
Long-running agents usually degrade as their histories grow: they repeat failed experiments, duplicate each other's work, or keep spending tokens without making progress. This paper's fix is structural rather than cosmetic - separate what the system remembers from what each agent sees, instead of trusting a bigger context window to sort it out.
Most agent benchmarks run short enough that this failure mode never surfaces, which says less about how good today's research agents are than about how little anyone has stress-tested them.