A new AI agent architecture called FlowState swaps full conversation history for structured memory of what the agent actually did, and gets better results for less money.
Researchers behind FlowState built a system that treats an agent's execution state, not its raw conversation history, as memory. Instead of storing everything verbatim, it keeps semantically typed state nodes, the relationships between them, and references to the original tool outputs. Two mechanisms do the work: Incremental State Update keeps the current state current as new information and feedback arrive, while Progressive State Access surfaces older state and supporting evidence only when the agent's reasoning actually needs it. Tested against a full-context baseline running the same DeepSeek-V4-Flash model, FlowState raised the average success rate on the MemoryArena benchmark by 4.55 percentage points and the pass rate on tau3-Bench by 13.95 percentage points, while cutting total token use by 43.2% and 40.6% on those two benchmarks respectively.
Context windows are the biggest cost lever for long-running AI agents, and most fixes so far have meant either paying for more tokens or risking that compression throws away details needed later. FlowState's bet is that an agent doesn't need to remember everything it said, just what it did, which is a smaller and more stable record to keep. If that holds up outside benchmark conditions, it points toward agents that get cheaper as tasks get longer, the opposite of how context costs usually behave.
It is also, notably, a paper revision rather than a first release, so treat the numbers as a claim still working through peer scrutiny, not a settled result.