AI/ ai · code-generation · plagiarism-detection · open-source-licensing

A Faster Way to Catch AI Code Plagiarism

A new hybrid search system traces AI-generated code snippets back to their original training examples without scanning billions of lines one by one.

A new search method can trace AI-written code snippets back to the exact training examples they came from, without scanning the entire training set each time.

Researchers built SourceTracker, a 300-million-parameter encoder for code retrieval, and paired it with a two-stage pipeline called HybridSourceTracker (HST). HST first uses vector search to narrow a massive corpus down to a small set of candidate snippets, then re-ranks those candidates with Winnowing, a classical code-fingerprinting technique, to confirm exact matches. The team trained and evaluated the system on a 10M-snippet subset of TheStackV2, a dataset used to train code-generating LLMs, including snippets with renamed identifiers meant to mimic how people tweak copied code. In a separate in-vitro test using a smaller 100k-snippet search space, HST matched Winnowing's accuracy on 30-token fragments and beat it by up to 5.4% once fragments reached 60 tokens or longer, all while keeping query time logarithmic instead of linear.

Winnowing-style fingerprint matching is accurate, but it checks a snippet against literally everything in the training set, which falls apart at the scale of modern code-LLM training data. HST's vector-search prefilter makes provenance checking plausible for billion-snippet corpora instead of just research-sized ones. That matters as more lawsuits and license disputes turn on whether a model's output counts as original work or a reproduction.

The catch: this only works if you can actually see what a model trained on, and most commercial code LLMs still will not disclose that, so the harder problem here is access, not math.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →