AI agents can now patch a crashing program while it is still running - and researchers just built the first real-world test of whether that is safe to trust.
A new paper introduces HealBench, a benchmark of 265 runtime errors pulled from 18 real-world code repositories, each paired with a reference fix from the actual patched version. The researchers also built HealGuard, a safety layer that forces healing code into an analyzable subset of Python and uses static and dynamic taint analysis to track whether a fix's changes reach code a developer explicitly protected. Testing a dedicated healing method plus three general-purpose coding agents across three backbone LLMs, the best setup resumed execution in 38.11% of cases and actually passed the target test in 28.68%. Among the runs that passed, HealGuard flagged 17.4% as having changes that could reach protected operations - and in controlled tests it caught every unsafe case, though at a 68.42% false-positive rate.
That gap between the program kept running and the fix was actually correct is the real story here. Prior work on this kind of self-healing code tested toy competition problems, not messy production repositories, so a one-in-three success rate on real code is a meaningful, sobering baseline rather than a victory lap.
A safety net that catches every real problem but cries wolf two-thirds of the time is still a safety net - just not one you can run unattended yet.