A new search algorithm for AI-driven scientific discovery would rather return nothing than return a lie.
Researchers behind a method called CISE, short for Conformal Interval-Driven Self-Evolution, are targeting a specific failure in AI systems that iteratively search for scientific breakthroughs. These systems often lean on cheap proxy reward signals to judge candidates, since running full, high-fidelity tests on every option is too expensive, and those proxies sometimes score infeasible candidates highly. CISE instead builds a statistically calibrated confidence interval for each candidate's properties, using conformal inference and density-ratio estimation, and only passes a candidate through if every required interval sits entirely inside the feasible zone. Tested on three self-evolving search tasks in materials science, every candidate CISE returned checked out under expensive, high-fidelity evaluation, while baseline methods returned longer lists that included candidates which failed the real test.
That distinction matters because validation, not idea generation, is the actual bottleneck in computational science: synthesizing and testing a candidate material in the real world costs far more than generating another guess. A shorter list of guaranteed winners is worth more to a lab with a limited budget than a longer list padded with maybes.
It is a quietly skeptical response to the current enthusiasm for AI that proposes thousands of candidates per hour: proposing fast is easy, and proving it right is still the hard part.