A new technique called UNBIND can make a code-generating AI model forget a specific chunk of memorized code without retraining it.
Researchers describe UNBIND as an inference-time method that steers a code large language model's hidden states away from reproducing a targeted snippet, rather than editing the model's weights. The system separates two problems: identifying which internal representations encode the target code, then building a separate direction to suppress its output. Tested across fourteen baseline methods, two code models, and two corpora, UNBIND cut exact reproduction of targeted code by 97.3 to 99.1 percent as measured by F-BLEU. In repeated extraction attempts, the number of snippets that could be pulled out as 50-token-plus exact matches dropped from as many as 262 out of 300 down to zero or two.
Code models memorize real functions from their training data, and that creates copyright and security exposure when the same code resurfaces verbatim in a generated completion. Retraining a model to delete that memory is slow and expensive, so a steering method applied only at inference time matters if it holds up outside benchmark conditions. UNBIND also preserved general coding ability well, solving at most two fewer HumanEval+ problems and six fewer MBPP+ problems than the unmodified models.
The paper's own numbers show the forgetting isn't total: even the best recovery attempts still pulled back up to 6.45 percent of a target snippet, a reminder that inference-time steering suppresses reproduction rather than erasing it from the model entirely.