Model Distillation Explained: Why Every AI Lab Ships It
Model distillation is the process of training a smaller "student" model to mimic a larger "teacher" model's outputs, and by mid-2026 every major AI lab ships a distilled sibling alongside its flagship. The pattern spans GPT-5.4 Mini, Gemini 3.1 Flash/Flash-Lite, Claude Haiku, and open releases like DeepSeek-R1-Distill. The bigger signal is that distillation has become the default architecture for shipping AI at production scale, not a discount option. For PMs, the real work is knowing when a distilled model quietly costs you accuracy versus when it's the obviously correct call.