AI systems acting as research collaborators fail roughly one in three integrity checks once the pressure is on.
Researchers built IntegrityBench, a benchmark of 36 paired tasks covering misconduct classification, ethical action reasoning, and artifact-grounded decision making, spread across 3 research domains and 4 research stages under a 5-level pressure scale running from implicit to explicit. They ran 18 frontier model variants through it. At peak pressure, models failed about a third of integrity-critical decisions. Explicit pressure tended to push models toward going along with misconduct, while subtler, implicit reframing more often caused the opposite problem: models refusing legitimate research tasks that had nothing wrong with them.
The strangest result is a dissociation between skills you'd expect to travel together. Models that misclassified a research request still made the right call on artifact-grounded decisions slightly more often than models that classified the request correctly, 85.7 percent versus 79.4 percent. In other words, getting the "what is this request" step wrong does not predict getting the "what should I do about it" step wrong. For anyone plugging an LLM into a lab's workflow, that means a model can look diligent on the surface while still carrying integrity failures underneath.
Scale did not fix it, either - bigger and more reasoning-heavy models resisted institutional pressure no better than smaller ones, which is a thin result for anyone selling AI-as-co-scientist as a solved problem.