Turns out you don't need a mountain of training prompts to teach a smaller AI model your best tricks.
An unreviewed preprint posted to arXiv, "Beyond Prompt Count: How Data Shapes Transfer in On-Policy Distillation," examines on-policy distillation, a technique where a smaller "student" model learns from a larger "teacher" model's feedback on the student's own generated answers. The researchers found that just four prompts from a dataset called DAPO produced roughly the same math score as 3,840 prompts from a different dataset, DeepMath. That efficiency was not a general rule, though: swapping which model served as teacher could flip which prompt set performed better, and a prompt source that failed on its own did not reliably help or hurt when mixed with others. The paper has not been peer reviewed.
Model distillation is one of the cheaper ways to shrink a large model into something faster and cheaper to run, and prompt selection has mostly been treated as an afterthought. This study argues the opposite: a handful of the right prompts can substitute for thousands, but only for that specific teacher-student pairing and task, meaning the shortcut does not automatically carry over to the next model you distill.
The paper's most deflating finding: careful prompt curation rarely beat plain random sampling. For an industry that loves to talk up its data curation pipelines, that is a fairly dry verdict.