Researchers have built an AI agent that is designed to doubt its own results before calling them discoveries.
The system, called EPOCH, targets a specific failure mode in AI research agents that search for programs, mathematical constructions, and proofs: they chase evaluator feedback without checking whether that feedback actually supports the claim being made. EPOCH's fix is procedural rather than algorithmic. It wraps the search in explicit task contracts, typed memory, active attempts to falsify its own candidates, admission checks, and independent replay, so a result has to survive scrutiny before it gets labeled a discovery. On the AlgoTune benchmark, EPOCH posted a mean normalized score of 0.65 against the best baseline's 0.53, and it also led on an internal Math14 suite and an AgentHPO aggregate.
The real news here isn't the score bump. It's the admission that most AI research agents have been grading their own homework. Optimizing for evaluator feedback without checking whether that feedback means what it looks like it means is exactly how a fragile, overfit candidate gets dressed up as a breakthrough. Building falsification and replay into the loop is a tacit acknowledgment that the current crop of discovery agents can't be trusted to self-report.
Whether this generalizes beyond AlgoTune and a homemade math suite is the open question - benchmarks designed in-house have a way of flattering the system built to beat them.