A new paper argues that most defenses against model distillation look better on paper than they actually are.
Researchers frame the problem as a game between a "teacher" model trying to stay useful while resisting copycats, and a "student" model trying to imitate it as cheaply as possible. They test existing defenses against an adaptive student that reweights the most valuable examples, rather than a passive one that just tries to copy everything equally. On math benchmarks GSM8K and MATH, that adaptive student recovers far more of the teacher's capability than passive testing suggested - a gap the authors call substantial. The team also proposes a cheap defense called Product-of-Experts, which blends the teacher's outputs with a proxy student's at generation time, using only a forward pass and no retraining.
This matters because providers often cite distillation defenses as evidence their models can't be cheaply cloned. If passive evaluation is the industry's default testing method, this suggests those robustness claims are inflated. Under adaptive testing, the gap between expensive defenses and the cheap PoE approach narrows a lot, while PoE stays far cheaper and keeps better reasoning quality.
The practical upshot is blunter than the marketing around antidistillation tools usually allows: strong distillation is still hard to stop, full stop. If cheap and expensive defenses perform similarly once you test them properly, providers spending heavily on defense may be buying a false sense of security rather than real protection.