AI/ ai-agents · llm-benchmarks · arxiv-research · agent-architecture

New Agent Architecture Splits Reasoning From Action Checks

A new arXiv study adds separate action-checking and completion-checking roles to AI agents, and the clearest gains go to weaker models, not top-tier ones.

DeReAct gives AI agents a built-in skeptic instead of letting one model grade its own homework.

Right now, most ReAct-style agents use a single language model to propose an action, carry it out, and decide when the task is done, which lets early mistakes snowball into false claims of success. DeReAct, described in a new arXiv paper, splits that job in two: a Critic checks a proposed action before it executes, and a Context Manager tracks what the environment has actually confirmed and certifies when a task is truly complete. Tested on the GAIA and SWE-bench Verified benchmarks, the setup lifted Pass@1 scores by 6.5 to 7.0 points for Qwen3-Coder-480B and by 4.2 to 5.2 points for Claude Sonnet 4.5.

Those numbers matter less than the pattern behind them: the benefit shrinks as the underlying model gets better. With Claude Opus 4.5, Pass@1 barely moved, though the agent still produced more evidence-backed trajectories with fewer unsupported completion claims. That suggests external oversight layers are compensating for weaknesses that stronger models have already learned to avoid on their own.

Worth remembering next time a benchmark post calls extra scaffolding a universal fix: sometimes it is just a patch for a model that was not that reliable to begin with.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →