AI/ ai agents · automated research · hallucination · benchmarks

New Agent Architecture Tries to Stop AI Research Papers From Lying

YouRA gives AI research agents persistent memory of evidence and failures, aiming to stop papers that claim more than experiments actually show.

Researchers have built an AI agent that keeps receipts.

A team describes YouRA ("Your Research Agent"), a new architecture meant to fix a specific flaw in autonomous research agents: the papers they write often claim more than the experiments they actually ran support. YouRA tracks hypotheses, gates, and evidence through what it calls a Verification State Architecture, pairs that with an Independent Controller that handles lifecycle and recovery decisions separately from execution, and logs failures as structured lessons through a Stateful Reflection module that can trigger repair, redesign, or a full reset. On MLR-Bench's ten-task end-to-end benchmark, YouRA beat two existing systems, MLR-Agent and AI Scientist V2, across all three model backbones tested. A diagnostic built on MLR-Bench's hallucination taxonomy also found fewer fact-based failures, and YouRA produced more outputs grounded in real data rather than invented results.

The real story here is less about beating a leaderboard and more about what the leaderboard is measuring. Fully autonomous research agents can now generate entire papers, methods sections and all, but if the claims don't trace back to code that actually ran, that is just confident-sounding fabrication with extra steps. YouRA's bet is that persistent, inspectable state, not a smarter language model, is what closes that gap, and the ablation results back that up: strip out either the verification state or the independent controller and performance drops below the full system.

That is a useful fix for a problem the field has mostly tried to paper over with better prompting. Whether "evidence-traceable" survives contact with messier, non-benchmark science is the question nobody's answered yet.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →