A new distillation technique teaches smaller AI models to care about what tokens mean, not just how likely they are.
Researchers have proposed WASD, short for Wasserstein-based knowledge distillation, a new way to train compact student language models to imitate larger teacher models, while actually accounting for which tokens mean similar things. Most existing distillation methods compare probability scores at each vocabulary slot without considering semantic closeness, so confusing one token for another costs the same whether they are synonyms or unrelated words. WASD fixes that by measuring the gap between teacher and student outputs with a Wasserstein distance built from a cost matrix derived from token embeddings, using a Sinkhorn divergence to keep training computationally tractable without adding extra neural networks. Across multiple model families and sizes, it improved performance on instruction following, math reasoning, and code generation.
Distillation is the standard trick for shipping chatbots and coding assistants that run fast and cheap without paying for a frontier-sized model at inference time, so any consistent accuracy bump compounds across every deployment that relies on it. The more interesting point is conceptual: token-embedding spaces already encode which words are close in meaning, and most distillation math has simply been ignoring that information, comparing probabilities as if every wrong token were equally wrong.
The code is open-source on GitHub, so the real test is whether the gains hold up outside the paper's own benchmarks, not just whether the math is elegant.