A team of researchers built a simulated academic world run entirely by AI agents, then used it to stress test how peer review breaks down.
The researchers introduced SciUtopia, a persistent, closed-loop simulation framework in which LLM agents play researchers, institutions, funders, and reviewers. The system models the full research lifecycle: choosing research directions, forming collaborations, submitting papers, peer reviewing, resubmitting after rejection, citing other work, securing funding, and eventually dropping out of research altogether. Across 61 separate simulated worlds, the team ran more than 40,000 simulated researchers across 8,000 institutions, producing roughly 400,000 publication decisions and 1.2 million LLM-written peer reviews. The code is public on GitHub.
The simulation surfaces dynamics that are hard to measure in the real publishing system. Papers rejected and resubmitted elsewhere pile up far more reviewing work than population growth alone would predict, and funding inequality between institutions can emerge even without any detectable compounding advantage from early grants. Both findings land squarely in an ongoing debate over whether peer review is already buckling under volume.
It is worth remembering this is a model built from models: LLM agents judging LLM-written papers inside an LLM-simulated economy, so the specific numbers describe a simulation's internal logic, not guaranteed facts about how real journals and grant committees behave.