A new multi-agent system runs the entire scientific method on AI agents, no human required to design the experiments.
Researchers built AEROBAT, a multi-agent system that automates behavioral scientific research on other AI agents. Given a target behavior, it generates hypotheses, designs controlled experiments, runs simulations, analyzes results, and writes up findings on its own. In testing, the researchers pointed AEROBAT at 12 target behaviors. It generated and tested 73 hypotheses, designed 1,160 controlled experiments, and executed 22,954 simulation rounds, turning up moderate-to-strong statistical evidence for 30 hypotheses, some of them novel.
The numbers matter more than the novelty. AI agents are already being deployed into messy, high-stakes environments, and understanding how they actually behave has been a slow, manual research bottleneck. A system that can run thousands of controlled trials in the time it takes a lab to design one could let behavioral scrutiny keep pace with deployment instead of trailing years behind it.
That said, automating the scientific method does not automate judgment. Thirty statistically supported hypotheses out of 73 tested is a research haul, not a verdict - someone still has to decide which findings are worth trusting and which are artifacts of how the experiments were designed.