A new framework lets deep-research AI agents build their plan as they go, instead of locking it in before they've seen enough evidence.
Researchers built DAGent as a directed acyclic graph system where an orchestrator expands the task graph one batch at a time, using confidence and uncertainty signals from finished steps to decide what to tackle next. That is a shift from the usual plan-then-patch approach, where agents sketch the whole plan upfront and fix it only after something breaks. A hierarchical context layer passes around compact summaries between agents by default, while keeping full execution traces on hand for recall. The team also built DAGRPO, a reinforcement-learning method that credits executor agents based on where they sit in the graph, not just whether the final answer was right.
On BrowseComp-Plus, GAIA, and xbench-DeepSearch, DAGent beat the strongest open-source baseline by 5.3, 5.8, and 2.0 points respectively at the Qwen3-235B-A22B scale, with the lead holding across four open backbones and extending to GPT-5 at a 327K-token context window. The bigger finding is efficiency: the same architecture reaches higher accuracy using fewer tokens, tool calls, and steps once planning is tied to evidence instead of committed upfront and patched later.
Benchmarks are not the same as surviving a messy real research task, so call this a strong lab result rather than proof deep-research agents are solved.