AI/ ai · llm-agents · reproducibility · social-science

AI Agents Can Reproduce Social Science Results From Methods Alone

A new paper tests whether AI agents can rebuild published social-science findings using only written methods sections, and most of the time they can.

AI agents can now rebuild a social-science paper's results from nothing but its methods section and original data.

The paper, titled "Read the Paper, Write the Code: Agentic Reproduction of Social-Science Results" (arXiv:2604.21965v2), describes an agentic reproduction system that extracts a structured methods description from a paper, then has agents reimplement the analysis under strict information isolation - they never see the original code, results, or the paper itself. A deterministic, cell-level comparison checks the reproduced output against the published numbers, and a separate error-attribution step traces any mismatch back to its source. The team ran four agent scaffolds and four LLMs across 48 papers with human-verified reproducibility. The source listing does not name the authors or their institution, so readers who want to verify the claim should check the arXiv page directly.

Results varied a lot by model, scaffold, and paper, but agents largely recovered the published findings. That is the real headline: reproducibility failures aren't only about missing code or data, they're often baked into how clearly a paper describes its own method. The root-cause analysis found failures split between agent errors and genuine underspecification in the papers themselves - meaning some failed reproductions were really the original authors' fault.

That's the uncomfortable part: a peer-reviewed methods section is supposed to be precise enough to rebuild the experiment, and if an AI agent can't parse it, a human replicator probably couldn't either.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →