Skip to content
estudIA

AI glossary

Distillation

Training a smaller “student” model to imitate a larger “teacher” model, keeping much of its quality at a fraction of the cost.

The large model produces answers (or probabilities) for many inputs, and the small model is trained to reproduce them. Because the teacher’s outputs are richer than plain right-or-wrong labels, the student learns faster than it would from scratch.

Distillation is one reason cheap, fast models keep getting better: the small tiers of a family often learn from the big one. Many providers’ terms forbid using their outputs to distil competing models.

Example: A company uses its large model to answer thousands of support questions and trains a small model on those answers. The small one replies almost as well on that topic, much faster and at a fraction of the cost.

In practice

  • If a small model surprises you with how well it works, it probably learned from a large one in its family.
  • Before distilling from a commercial model’s answers, read its terms of use.
  • A distilled model usually does well at what it was taught and worse outside that area.

Related terms

← Back to the glossary