Skip to content
Chatbotscape
The papers, Hugging Face documentation, OpenAI's distillation announcement and the Anthropic and OpenAI terms were read on 9 October 2026. Provider workflows and terms change, so check the current version before you build on them.
Model Distillation· LLM cost and latency
Model distillation trains a smaller model (the student) to reproduce the behavior of a larger model (the teacher) on a task. The student learns from what the teacher outputs, so it can be cheaper and faster to run than the teacher, within the limits of what it was shown.
By Chatbotscape Editorial· Methodology· Published 10 October 2026· Updated 10 October 2026

Model Distillation — Teaching a Small Model to Imitate a Big One

Quick answer: A large language model answers well but costs more and takes longer per reply. In distillation you let that large model (the teacher) answer a set of questions, then train a small model (the student) to produce similar answers. If it works, you run the small model for that task and pay less per answer. The catch is that the student only copies what it saw. Questions unlike the training set can get noticeably worse answers, so a distilled model has to be tested against the teacher before it takes over.

How distillation works

Geoffrey Hinton, Oriol Vinyals and Jeff Dean gave the technique its name in a 2015 paper, "Distilling the Knowledge in a Neural Network." Its starting point was practical: averaging the predictions of many models improves results, but a group of models is expensive to deploy, so the authors looked for a way to compress that knowledge into one model. Today the vocabulary is simple:

  • Teacher. The large, capable model whose behavior you want to copy.
  • Student. The small model you train to imitate it.
  • Training set. Questions or inputs, together with the teacher's answers. Quality here decides the result.

The Hugging Face documentation shows the classic version in code. The teacher's raw scores are softened with a temperature setting, so the student learns not only the top answer but how the teacher weighed the alternatives, and the student's loss combines two parts: how closely it matches the softened teacher, and how well it does on the true labels. That is a different "temperature" from the one that sets how random a chatbot's wording is.

Two forms you will meet

Soft-label distillationOutput-based distillation
What the student learns fromThe teacher's internal scores (logits) for each inputThe teacher's finished text answers
What you needAccess to the teacher's internalsOnly the teacher's replies
Typical settingOpen models you can run yourselfHosted models reached through an API
ExampleHugging Face image-classification walkthroughOpenAI's stored-completions workflow

We infer, and have not tested, that most businesses using a hosted model through an API only have the second option, because such APIs usually return text rather than internal scores. Check your provider's documentation.

OpenAI's 1 October 2024 announcement of "Model Distillation in the API" shows the output-based version end to end. You create an evaluation for the target task, store the larger model's real outputs, review and filter them into a training set, fine-tune the smaller model on it, then test the result in the evaluation against the larger model and repeat if it falls short. The post says stored completions are free to use and that training and running a distilled model cost the same as standard fine-tuning. Treat those details as dated and read the current documentation.

How distillation differs from similar ideas

TechniqueWhat changesNeeds a teacher model?
DistillationA smaller model is trained to imitate a larger oneYes
Fine-tuningAn existing model is trained further on your examplesNo, though the examples can come from a teacher
QuantizationThe same model's numbers are stored with less precision, making it smallerNo
Retrieval (RAG)Facts are fetched at answer time; no model is trainedNo

Distillation and fine-tuning overlap: the output-based workflow above is fine-tuning a small model on a large model's answers. The useful distinction is the purpose. Fine-tuning teaches a model your task; distillation specifically aims to move a task from an expensive model to a cheap one. And neither is a way to add up-to-date facts. For that, see fine-tuning versus RAG.

What it can do for a chatbot

The case for distillation is cost and speed on a narrow, repetitive task, such as classifying incoming messages or answering a fixed set of question types. The size of the gain depends on the models, so measure your own. The best-known published example is DistilBERT, whose authors report a BERT model made 40 percent smaller that keeps 97 percent of its language-understanding performance and runs 60 percent faster. That is one team's result for one older model on specific benchmarks, not a promise for a modern chatbot.

Savings arrive as fewer tokens billed at a lower rate per answer, which shows up in cost per conversation. Other levers are often cheaper to try first: semantic caching, prompt caching and shorter prompts, all in the reduce chatbot costs guide.

Limits and risks

  • The student inherits the teacher's mistakes. If the teacher produced an AI hallucination that made it into the training set, the student learns it.
  • Coverage is narrow. The student is good on questions that resemble its training data and can be weak elsewhere. A customer who asks something new meets the weak spot.
  • Safety behavior is not guaranteed to carry over. Refusals, tone rules and handoff behavior should be tested again on the student, together with your AI guardrails. We found no source that promises they transfer.
  • It is a maintenance job. When your products or policies change, the training set and the student go out of date, a form of model drift.
  • It needs a test. Without a golden dataset and a way to grade answers, such as LLM-as-a-judge, you cannot tell whether the student is good enough.

Check the terms before you distill from an API

Some providers restrict using their output to build other models. Anthropic's Commercial Terms, section D.4, say a customer may not "access the Services to build a competing product or service, including to train competing AI models." OpenAI's Terms of Use list among things you cannot do: "Use Output to develop models that compete with OpenAI." Training a small model for your own chatbot may or may not fall under such clauses, and that is a question for the provider's current terms and, if needed, a lawyer, not for this page. We are not lawyers and this is not legal advice.

FAQ

What is model distillation?

Training a small model to imitate a large one on a task, so the small model can be used instead and cost less per answer.

Is distillation the same as fine-tuning?

They overlap. Distilling from API outputs is done by fine-tuning a small model on a large model's answers. The aim differs: fine-tuning adapts a model to a task, distillation moves a task from an expensive model to a cheap one.

Does a distilled model give the same answers as the original?

No. It approximates the teacher on the kinds of input it was trained on and can be worse elsewhere. Test it on your own questions.

Do I need to distill a model to use a cheaper one?

No. You can often switch to a smaller existing model and test it directly. The guide to testing a smaller model shows how.

Is it allowed to train a model on another model's outputs?

It depends on the provider's terms. Anthropic and OpenAI both have clauses about building competing models; read the current terms before you start.

Sources