Model Distillation· LLM cost and latency
Model Distillation — Teaching a Small Model to Imitate a Big One
Quick answer: A large language model answers well but costs more and takes longer per reply. In distillation you let that large model (the teacher) answer a set of questions, then train a small model (the student) to produce similar answers. If it works, you run the small model for that task and pay less per answer. The catch is that the student only copies what it saw. Questions unlike the training set can get noticeably worse answers, so a distilled model has to be tested against the teacher before it takes over.
How distillation works
Geoffrey Hinton, Oriol Vinyals and Jeff Dean gave the technique its name in a 2015 paper, "Distilling the Knowledge in a Neural Network." Its starting point was practical: averaging the predictions of many models improves results, but a group of models is expensive to deploy, so the authors looked for a way to compress that knowledge into one model. Today the vocabulary is simple:
- Teacher. The large, capable model whose behavior you want to copy.
- Student. The small model you train to imitate it.
- Training set. Questions or inputs, together with the teacher's answers. Quality here decides the result.
The Hugging Face documentation shows the classic version in code. The teacher's raw scores are softened with a temperature setting, so the student learns not only the top answer but how the teacher weighed the alternatives, and the student's loss combines two parts: how closely it matches the softened teacher, and how well it does on the true labels. That is a different "temperature" from the one that sets how random a chatbot's wording is.
Two forms you will meet
| Soft-label distillation | Output-based distillation | |
|---|---|---|
| What the student learns from | The teacher's internal scores (logits) for each input | The teacher's finished text answers |
| What you need | Access to the teacher's internals | Only the teacher's replies |
| Typical setting | Open models you can run yourself | Hosted models reached through an API |
| Example | Hugging Face image-classification walkthrough | OpenAI's stored-completions workflow |
We infer, and have not tested, that most businesses using a hosted model through an API only have the second option, because such APIs usually return text rather than internal scores. Check your provider's documentation.
OpenAI's 1 October 2024 announcement of "Model Distillation in the API" shows the output-based version end to end. You create an evaluation for the target task, store the larger model's real outputs, review and filter them into a training set, fine-tune the smaller model on it, then test the result in the evaluation against the larger model and repeat if it falls short. The post says stored completions are free to use and that training and running a distilled model cost the same as standard fine-tuning. Treat those details as dated and read the current documentation.
How distillation differs from similar ideas
| Technique | What changes | Needs a teacher model? |
|---|---|---|
| Distillation | A smaller model is trained to imitate a larger one | Yes |
| Fine-tuning | An existing model is trained further on your examples | No, though the examples can come from a teacher |
| Quantization | The same model's numbers are stored with less precision, making it smaller | No |
| Retrieval (RAG) | Facts are fetched at answer time; no model is trained | No |
Distillation and fine-tuning overlap: the output-based workflow above is fine-tuning a small model on a large model's answers. The useful distinction is the purpose. Fine-tuning teaches a model your task; distillation specifically aims to move a task from an expensive model to a cheap one. And neither is a way to add up-to-date facts. For that, see fine-tuning versus RAG.
What it can do for a chatbot
The case for distillation is cost and speed on a narrow, repetitive task, such as classifying incoming messages or answering a fixed set of question types. The size of the gain depends on the models, so measure your own. The best-known published example is DistilBERT, whose authors report a BERT model made 40 percent smaller that keeps 97 percent of its language-understanding performance and runs 60 percent faster. That is one team's result for one older model on specific benchmarks, not a promise for a modern chatbot.
Savings arrive as fewer tokens billed at a lower rate per answer, which shows up in cost per conversation. Other levers are often cheaper to try first: semantic caching, prompt caching and shorter prompts, all in the reduce chatbot costs guide.
Limits and risks
- The student inherits the teacher's mistakes. If the teacher produced an AI hallucination that made it into the training set, the student learns it.
- Coverage is narrow. The student is good on questions that resemble its training data and can be weak elsewhere. A customer who asks something new meets the weak spot.
- Safety behavior is not guaranteed to carry over. Refusals, tone rules and handoff behavior should be tested again on the student, together with your AI guardrails. We found no source that promises they transfer.
- It is a maintenance job. When your products or policies change, the training set and the student go out of date, a form of model drift.
- It needs a test. Without a golden dataset and a way to grade answers, such as LLM-as-a-judge, you cannot tell whether the student is good enough.
Check the terms before you distill from an API
Some providers restrict using their output to build other models. Anthropic's Commercial Terms, section D.4, say a customer may not "access the Services to build a competing product or service, including to train competing AI models." OpenAI's Terms of Use list among things you cannot do: "Use Output to develop models that compete with OpenAI." Training a small model for your own chatbot may or may not fall under such clauses, and that is a question for the provider's current terms and, if needed, a lawyer, not for this page. We are not lawyers and this is not legal advice.
FAQ
What is model distillation?
Training a small model to imitate a large one on a task, so the small model can be used instead and cost less per answer.
Is distillation the same as fine-tuning?
They overlap. Distilling from API outputs is done by fine-tuning a small model on a large model's answers. The aim differs: fine-tuning adapts a model to a task, distillation moves a task from an expensive model to a cheap one.
Does a distilled model give the same answers as the original?
No. It approximates the teacher on the kinds of input it was trained on and can be worse elsewhere. Test it on your own questions.
Do I need to distill a model to use a cheaper one?
No. You can often switch to a smaller existing model and test it directly. The guide to testing a smaller model shows how.
Is it allowed to train a model on another model's outputs?
It depends on the provider's terms. Anthropic and OpenAI both have clauses about building competing models; read the current terms before you start.
Related terms
- Fine-tuning — the training step that output-based distillation uses.
- Large language model — what the teacher usually is.
- LLM-as-a-judge — a way to grade the student against the teacher.
- Golden dataset — the fixed questions you test with.
- LLM token — the unit a cheaper model bills less for.
- AI hallucination — the error a student can inherit.
Sources
- Hinton, Vinyals and Dean, Distilling the Knowledge in a Neural Network — arxiv.org/abs/1503.02531, read 9 October 2026. Source of the origin of the term and the ensemble-compression motive.
- Sanh et al., DistilBERT, a distilled version of BERT — arxiv.org/abs/1910.01108, read 9 October 2026. Source of the 40 percent, 97 percent and 60 percent figures, which are the authors' results.
- Hugging Face, Knowledge Distillation for Image Classification — huggingface.co/docs/transformers/en/tasks/knowledge_distillation_for_image_classification, read 9 October 2026. Source of the teacher, student, temperature and two-part loss description.
- OpenAI, Model Distillation in the API — openai.com/index/api-model-distillation/, published 1 October 2024, read 9 October 2026. Source of the stored-completions, evals and fine-tuning workflow.
- Anthropic, Commercial Terms of Service — anthropic.com/legal/commercial-terms, section D.4, read 9 October 2026.
- OpenAI, Terms of Use — openai.com/policies/terms-of-use, read 9 October 2026.
- Chatbotscape evaluation methodology. /methodology (continuously updated).