Skip to content
Chatbotscape
Editorial flat-vector illustration for How to Test Whether a Smaller, Cheaper AI Model Is Good Enough for Your Chatbot
Published 8 min read

How to Test Whether a Smaller, Cheaper AI Model Is Good Enough for Your Chatbot

Quick answer: Do not decide from a vendor chart. Build a test set of 150 to 300 real customer questions, run the current model and the candidate on every one, grade both with a model-based judge, check a sample by hand, and set a pass bar before you look at the results. If the smaller model passes on easy questions but fails on hard ones, send only the easy ones to it. In our invented example a 70/30 split saves about $170 a month; the test itself costs under $10.

Step 1: Build the test set from real questions

Pull a month of real conversations (anonymized; see the data privacy guide). Sample 150 to 300 customer messages across your main topics, and deliberately include the hard ones: multi-turn follow-ups, angry customers, questions the bot should refuse, and cases where it should hand off to a person. Write down what a good answer must contain for each, or keep the answer you have already approved. This is a golden dataset; it is also the asset you reuse in the regression testing guide.

Weight it like your traffic, then add extra hard cases. A set made only of easy questions will pass every model.

Step 2: Choose the candidate and freeze everything else

Pick one smaller model, often from the same provider so the API and behavior are close. Keep the prompt, the knowledge base, the retrieval settings and the generation settings identical for both runs. If you change the prompt for the candidate, you are testing the prompt, not the model. Note that a smaller model sometimes needs a tighter prompt; if so, run the comparison twice and report both.

For context, the best-known distillation result, DistilBERT, kept 97 percent of its parent's language-understanding score at 40 percent smaller (the authors' figures on BERT benchmarks). Do not expect a number like that on your chatbot. Your questions decide the outcome.

Step 3: Run both models on every question

Run the full set through the current model and the candidate. For multi-turn cases, replay the same earlier turns for both. Save every answer, the tokens used and the response time. Run each question a few times if your temperature setting is above zero, since answers vary.

Step 4: Grade with a judge, then check by hand

Grade each pair with LLM-as-a-judge: a separate model scores each answer against your criteria, such as correct facts, follows policy, right tone, hands off when it should. Use a judge that is not the candidate, show the two answers in both orders, and treat a verdict that changes when the order is swapped as a tie. These precautions address the known biases described in the glossary entry.

Then read a sample yourself. Our rule: read at least 40 pairs, every pair where the judge and the cheaper model disagree with the current model, and every case that involved a refusal or handoff. A judge can miss tone problems and subtle factual errors, and it can prefer a longer answer. For the retrieval side of a RAG bot, add the checks in the RAG evaluation guide.

Step 5: Set the pass bar before you look

Write the bar down first. Ours, which you should adjust, has four lines:

  1. Quality. The candidate is judged as good as or better than the current model on at least a set share of questions that you choose, for example 90 percent.
  2. No new serious errors. Zero wrong answers about prices, policies or safety that the current model got right.
  3. Behavior. Refusals and handoffs trigger as often as before, with your AI guardrails intact. See the handoff design guide.
  4. Speed. Response time is no worse than before for the traffic you move.

Splitting by topic matters. A candidate that fails overall may still pass on order-status lookups or hours-and-location questions, and that is the basis for Step 7.

Step 6: Check the arithmetic

Do the sum with your own token counts, using the method in the token usage and cost estimate guide. Every number below is invented for illustration. They are not vendor prices.

ItemValue
Conversations per month30,000
Tokens per conversation1,500 in, 300 out
Large model, invented price$3 in and $15 out per million tokens
Small model, invented price$0.30 in and $1.50 out per million tokens
Large model cost per conversation$0.0045 + $0.0045 = $0.009
Small model cost per conversation$0.00045 + $0.00045 = $0.0009
All on large30,000 × $0.009 = $270.00
All on small30,000 × $0.0009 = $27.00 (saving $243.00, if it passed everywhere)
70% to small, 30% to large21,000 × $0.0009 + 9,000 × $0.009 = $18.90 + $81.00 = $99.90
Net saving from the split$270.00 − $99.90 = $170.10 a month
One-time test cost300 questions × ($0.009 + $0.0009) ≈ $2.97, plus 600 judge calls at an assumed $0.01 = $6.00, total about $9

Two lessons from the invented table. The test is cheap compared with the saving, so there is little reason to skip it. And the split gets 70 percent of the saving without exposing hard questions to the weaker model, though it adds the work of the router in Step 7. If the saving at your real volume is a few dollars a month, stop here: a model swap is not worth the risk. Try cheaper levers first, listed in the reduce chatbot costs guide.

Step 7: Roll out gradually, with a router if needed

Do not switch everyone at once.

  • Route by topic or difficulty. A router sends easy, well-tested question types to the small model and everything else to the large one. A router can be a simple rule (topic labels from your own classifier) or a small model that decides; start with rules, because they are easier to explain and debug.
  • Start with a slice. Send 5 to 10 percent of eligible conversations to the candidate for a week, compare satisfaction, handoff rate and complaints with the rest, then widen.
  • Keep a way back. Make the model a setting, not a code change, so you can revert in minutes.
  • Keep watching. Add the candidate to your production monitoring and rerun the test set when the provider updates the model or when you change the prompt.

Where the task is narrow and high volume, and a smaller model nearly passes, distillation or fine-tuning on approved answers may close the gap. That is a bigger project; read the fine-tune versus RAG guide first, and check the provider's terms.

Common mistakes

  • Testing on easy questions. Include refusals, follow-ups and angry customers.
  • Changing the prompt between runs and then crediting or blaming the model.
  • Trusting the judge alone. Read a sample by hand.
  • Setting the bar after seeing the results, which tunes it to the answer you wanted.
  • Switching everyone at once with no way back.
  • Skipping the arithmetic, then taking on risk for a saving that is too small to notice.

FAQ

Can I switch my support chatbot to a smaller model without losing quality?

Sometimes, for some question types. Run both models on a golden set of your own questions, grade them, and move only the topics where the smaller model passes.

How many test questions do I need?

We use 150 to 300 with extra hard cases. That is enough to see large differences, not small ones; treat results near the bar with caution.

What is the difference between a smaller model and a distilled model?

A distilled model is a small model trained to imitate a larger one on a task. Any small model can be tested with this method; see model distillation.

What is LLM routing?

Sending each request to the model that suits it, such as a small model for simple questions and a large one for complex ones, based on rules or a classifier.

How do I know the judge is fair?

Swap answer order, use a judge different from the candidate, and compare its verdicts with your own on a sample.

When should I not switch?

When the saving is small, when the candidate fails refusal or handoff checks, or when you cannot test properly.

Sources

  • OpenAI, Model Distillation in the API — openai.com/index/api-model-distillation/, published 1 October 2024, read 9 October 2026: the loop of creating an eval, collecting outputs, fine-tuning a smaller model and testing it against the larger one.
  • Sanh et al., DistilBERT, a distilled version of BERT — arxiv.org/abs/1910.01108, read 9 October 2026: 40 percent smaller, 97 percent of language understanding retained, 60 percent faster, the authors' results.
  • Chatbotscape evaluation methodology. /methodology (continuously updated).

About this guide

Chatbotscape launched in 2026 as an independent review site for chatbot platforms. This guide is editorial guidance; the seven-step test, the pass bar and the worked example are our own practice with invented numbers, and we have not benchmarked any model.

Last updated

10 October 2026 — first published.