A new search method can trace AI-written code snippets back to the exact training examples they came from, without scanning the entire training set each time.
Researchers built SourceTracker, a 300-million-parameter encoder for code retrieval, and paired it with a two-stage pipeline called HybridSourceTracker (HST). HST first uses vector search to narrow a massive corpus down to a small set of candidate snippets, then re-ranks those candidates with Winnowing, a classical code-fingerprinting technique, to confirm exact matches. The team trained and evaluated the system on a 10M-snippet subset of TheStackV2, a dataset used to train code-generating LLMs, including snippets with renamed identifiers meant to mimic how people tweak copied code. In a separate in-vitro test using a smaller 100k-snippet search space, HST matched Winnowing's accuracy on 30-token fragments and beat it by up to 5.4% once fragments reached 60 tokens or longer, all while keeping query time logarithmic instead of linear.
Winnowing-style fingerprint matching is accurate, but it checks a snippet against literally everything in the training set, which falls apart at the scale of modern code-LLM training data. HST's vector-search prefilter makes provenance checking plausible for billion-snippet corpora instead of just research-sized ones. That matters as more lawsuits and license disputes turn on whether a model's output counts as original work or a reproduction.
The catch: this only works if you can actually see what a model trained on, and most commercial code LLMs still will not disclose that, so the harder problem here is access, not math.