AI/ ai · ai-agents · agentevolver · benchmarks

AgentEvolver Lets AI Agents Upgrade Themselves Mid-Task

AgentEvolver captures what AI agents learn mid-task into reusable skills, posting an 82.08% SWE-bench Pro Public score without retraining the model.

A new research system lets AI agents rewrite their own tools and habits while they work, without touching the underlying model.

Researchers released AgentEvolver, a framework that captures what an AI agent learns while finishing a task and turns it into reusable capability: new operations, methods, sub-agents, control flow, interfaces, and supporting state, instead of letting every lesson evaporate after one conversation. A shared Runtime tracks these components through a common versioned lifecycle, while persistent planning and recoverable context preserve the agent's goals and supporting evidence across runs. On the SWE-bench Pro Public benchmark, the team reports an 82.08% resolution rate with evolution enabled, beating its own baseline without it. The researchers also tested the system on six other applications, including website, game, and research tasks, and found that some capabilities carried over successfully while other attempts stalled or failed outright.

Most agentic AI demos show a model nailing one task, then forgetting everything by the next session; AgentEvolver's bet is that the surrounding system, not the model weights, can accumulate experience, a cheaper and more auditable path to improvement than fine-tuning. That framing also forces a distinction leaderboard scores usually blur: whether an agent got lucky once or actually got better at the underlying skill.

The paper's most refreshing move is admitting its own gaps, logging a failed strategy alongside the wins and leaving transfer to new tasks and total development cost as open questions.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →