Researchers have built an AI system that studies its own mistakes to write better scientific formulas.
Symbolic regression is the search for compact mathematical equations that fit a dataset, an alternative to opaque neural-network models. A new method called RISR changes how large language models tackle that search: instead of judging a candidate formula only by its overall fit score, RISR encodes the pattern of leftover errors - the residuals - across every data point and feeds that information back into the model as it proposes revisions. A second component then predicts whether a candidate correction is actually worth fitting before the system commits computing power to it. On the LLM-SRBench benchmark, RISR hit 63.57% accuracy within a 1% error tolerance and 38.50% within a stricter 0.1% tolerance on in-distribution problems, with 56.07% and 38.24% on out-of-distribution problems, beating baselines that use the same underlying language model.
Most LLM-based symbolic regression tools treat a good-enough fit as success and discard information about where a formula goes wrong. RISR's residual-reading approach lets it target the specific part of an equation that is failing, a more surgical fix than the usual generate-and-rescore loop. That matters for fields like physics and biology, where a formula that is accurate everywhere except one regime is often useless.
Still, these are modest gains, not a breakthrough: even the best in-distribution score tops out below two-thirds accuracy at a loose 1% tolerance, and it drops below 40% once the tolerance tightens. Reading your own errors helps, but it does not yet solve equation discovery.