An AI coding agent spent five weeks improving the math behind DNA barcodes, nudging one key bound from 114 to 120.
Researchers set an LLM-based coding agent loose on an open problem in coding theory: finding the largest possible set of four-letter sequences, the kind used as DNA barcodes, that all stay a safe number of edits apart from each other. The agent wrote its own search code and verification scripts; a human picked the problem and designed the checking protocol. By restricting the search to codes with a specific symmetry, a known shortcut that shrinks the search space about fourfold, the agent raised the best known bound for length-6, distance-3 codes from 114 words to 120. The same pipeline improved twelve other lower bounds for codes of length 6 to 9 and distance 3 to 6.
This fits a pattern taking shape across math research, where agents like Google DeepMind's FunSearch and AlphaEvolve have shown LLMs can grind out small, verifiable improvements to standing bounds faster than humans alone. The more interesting finding here is methodological, not numerical: the team's own early search stalled at 116, and a second run with a better-tuned search operator is what actually found 120, meaning the gain came from search engineering as much as from the model itself.
The paper is just as candid about its mistakes. A later claim that the method failed to generalize to length 7 turned out to be wrong, undone by the same bug: an intermediate result that got written down, never rechecked, and treated as settled fact. A separate instance of that exact error had already cost the team three weeks elsewhere. Checking only final outputs, as their protocol required, never caught it - for a paper about error-correcting codes, that's a tidy irony.