Distillation trains a smaller "student" model to imitate a larger "teacher" model.
How It Works
- Generate outputs from the teacher on many inputs.
- Train the student to reproduce them — either final answers or full probability distributions over tokens, which carry richer information.
Benefits
- Much of the teacher's quality on target tasks at a fraction of the cost.
- Faster, cheaper inference.
- Models small enough for devices or high-volume use.
Task-Specific Distillation
For a specific application — say, classifying support tickets — distilling from a large model onto a small one trained just for that task can be very effective.
Data Matters
The student learns only what the teacher demonstrates. Cover the variety of inputs expected in production, including edge cases.
Limitations
- The student rarely matches the teacher on broad, complex tasks.
- Teacher errors are copied.
Terms of Use
Many model providers restrict using their outputs to train competing models. Check terms before distilling from a commercial API.