Small AI models just got a cheat sheet written by a bigger model, and it beats the usual tricks for teaching them specialized knowledge.
Researchers built a technique called instruction retrieval. A teacher model looks at a domain's problems, groups them into clusters, and for each cluster writes one instruction containing the background knowledge the cluster depends on, a procedure for solving that kind of problem, and the common mistakes made on it. That corpus only needs to be built once per domain, and the teacher is not needed again afterward. At inference time, a frozen small model retrieves the instructions closest to its question and follows them, with no fine-tuning. Tested across medicine, law, and math benchmarks, the approach beat zero-shot prompting on every task and beat few-shot prompting, self-consistency, and plain text retrieval on medicine and law. On the MedQA benchmark, the instruction corpus raised accuracy by 10.6 points, compared with 3.9 points from retrieving textbook passages, and few-shot examples from the same teacher actually made results worse.
Small models run on phones, laptops, and other edge devices because they are cheap and keep data local, but they stumble on expert questions that need deep domain knowledge. Fine-tuning can fix that, but it has to be redone for every model and every domain. Dumping retrieved text, like a textbook paragraph, just hands the model raw material and trusts it to figure out what to do with it. This method split facts from procedure, and the researchers' error analysis backs that split up: background knowledge fixed questions the small model always got wrong, while the procedural tips and mistake warnings only helped on questions where it was already wavering between answers.
It is not fine-tuning and it is not magic. It is a well-organized cheat sheet, and the paper's case is that organizing it well beats dumping more raw text on a small model.