
Few-Shot Examples for a Support Chatbot
How Many to Write, Which Ones, and How to Test Them
Quick answer: Start with the instruction and add examples only where the bot gets something wrong. Most vendors point to a small number: Anthropic recommends 3 to 5, and Google warns that too many can make the model overfit to them. Choose examples that sit on your hardest cases, not your easiest, keep their format identical, vary their wording so the bot does not memorize one phrasing, and prove they helped by running the same 30 real messages with and without them. Examples fix format, tone and category boundaries. They do not fix missing facts.
Step 1: Decide whether you need examples at all
Examples cost something on every call, so earn them. Write the instruction first, run 30 real customer messages through it, and read the failures. Examples are the right fix when the failures are about the shape of the answer: the bot is too long, uses the wrong greeting, ignores your required layout, or labels two neighboring categories inconsistently. They are the wrong fix when the failures are about facts. A bot that quotes last year's price needs a corrected knowledge base, not an example; see reduce chatbot hallucinations.
A quick sort of failures to the right lever:
| What goes wrong | Likely lever |
|---|---|
| Wrong facts, prices, policies | Knowledge base and retrieval |
| Wrong layout or length | Instruction first, then examples |
| Tone is off | Instruction, then two or three tone examples |
| Two categories confused | Rewrite the labels, then add boundary examples |
| Needs a tool or action | Tool definitions, not examples |
| Fails on rare phrasing at scale | Consider fine-tuning |
Step 2: Write three to five examples, and make them hard ones
Vendor advice agrees on the outline. Anthropic's prompting guide says to "include 3–5 examples for best results" and to make them relevant, diverse and structured, "cover edge cases and vary enough that Claude doesn't pick up unintended patterns." OpenAI's guide says to "show a diverse range of possible inputs with the desired outputs." Google's guide recommends always including examples while adding that "you may need to experiment with the number of examples."
Turn that into a choice rule. Pick your examples from the failures you found in Step 1, not from the clean cases the bot already handles. A useful spread for a support bot, which is our own suggestion rather than a published formula:
- One typical case, so the baseline layout is shown.
- One boundary case, a message that sits between two categories, answered the way you want it resolved.
- One refusal or out-of-scope case, showing exactly how the bot declines and what it offers next.
- One handoff case, showing the wording that passes a customer to a person, consistent with your escalation playbook.
- Optional: one messy case, with typos or two questions in one message.
Here is an invented pair, using a refund policy of 30 days purely as a placeholder:
<example>
Customer: can i still return the jacket i bought 5 weeks ago??
Reply: Our return window is {return_window} days from delivery, so a purchase from five weeks ago is outside it. I can pass your request to the support team, who can look at exceptions. Would you like me to do that?
</example>
Notice the placeholder. Facts that change should not live in the examples at all; put them in the knowledge base or the system prompt and let the example show only how the answer is phrased and structured.
Step 3: Keep the format identical, and mark examples as examples
Google's guide says: "Make sure that the structure and formatting of few-shot examples are the same to avoid responses with undesired formats." Treat that literally. Same labels, same punctuation, same line breaks, same length range in every example. One example with a bulleted reply teaches the model that bullets are allowed.
Mark the examples so the model cannot confuse them with instructions or with the live customer message. Anthropic recommends wrapping them in <example> tags and multiple examples in <examples> tags; OpenAI's guide demonstrates the same idea with tagged blocks and suggests placing examples in the developer message. Whichever platform you use, the principle is the same: a clear boundary between the instruction, the examples and the live message. This also reduces the chance that text a customer types inside a message is read as part of your prompt, a risk covered under prompt injection.
Step 4: Order, selection and the bot that repeats itself
Order. Research on older models found that order alone could swing accuracy: Zhao and colleagues (2021) reported results ranging "from near chance to near state-of-the-art," with a bias toward answers placed near the end of the prompt. Current models may be less fragile, but the cheap safe habit is to avoid putting all examples of one category together, mix the order, and re-test if you reorder.
Static or dynamic. A static set is the same examples every time, easy to review and cache. A dynamic set retrieves the examples closest to the incoming message. Liu and colleagues (2021) found similarity-based selection beat random selection on the benchmarks they tested. Dynamic selection needs embeddings and a store of approved examples, so it is a later upgrade, not a first step.
When the bot copies the examples. If customers get replies that echo your sample sentences, the cause is usually narrow examples. Google's guide warns that "if you include too many examples, the model may start to overfit the response to the examples," and Anthropic's says examples should vary enough that the model does not pick up unintended patterns. Our fixes, in order of effort: vary the sentence structure across examples; replace concrete details with placeholders like {return_window}; cut the number of examples to the minimum that fixes the failure; and state in the instruction that the examples show style, not wording to repeat.
Step 5: Test it before and after (our method)
This is Chatbotscape's own procedure, not a published standard.
- Take 30 real customer messages, including five out-of-scope ones, with a reference answer or correct category for each. Reuse your golden dataset if you have one.
- Run A: the instruction only. Run B: instruction plus your three to five examples. If you are tempted to add more, make that Run C and change nothing else.
- Score each run on three things: whether the answer is correct, whether the format matches your layout, and whether it copies an example.
- For copying, use a mechanical check: flag any reply that shares a run of six or more consecutive words with an example. The six-word threshold is our own rule of thumb.
- Keep the examples only if B beats A on correctness or format without raising copying. If C beats B by a small margin, prefer B, since every extra example is paid for on every message.
Do not test with the messages you used to write the examples, or you are grading the bot on its own answers. Hold out at least half. Then keep the set as a regression check: the chatbot regression testing guide shows how, and any model change can alter how examples are read, so re-run it.
Few-shot, RAG or fine-tuning?
They solve different problems. Examples change how the bot answers. Retrieval changes what it knows. Fine-tuning changes the model, at the price of a dataset and upkeep. OpenAI's guide presents few-shot learning as a way to steer a model toward a task without fine-tuning it. Our fine-tune vs RAG guide covers when the heavier options are worth it. Examples also share the context window with retrieved passages and conversation history, so a long example block can crowd out the facts the answer needs.
Common mistakes
Examples that all look alike. The bot learns the sameness. Vary length, phrasing and customer wording.
Facts inside examples. They go stale and the bot keeps repeating them. Use placeholders.
Only easy cases. The bot already handles those. Spend examples on the failures.
Examples that contradict the rules. If an example promises something the instruction forbids, the concrete example is the stronger signal.
Changing examples and temperature together. You will not know which helped; see the temperature settings guide.
Never retesting. A model upgrade can change how the same prompt behaves.
FAQ
How many few-shot examples should I use in a chatbot prompt?
Anthropic recommends 3 to 5. Google says to experiment and warns that too many can overfit. Start with three on your hardest cases and add one at a time only if your test shows a gain.
Why does my chatbot repeat my example answers word for word?
Usually the examples are too similar or too specific. Vary the wording, use placeholders for details, reduce the count, and tell the bot the examples show style rather than text to reuse.
Does the order of few-shot examples matter?
It did on older models, in some cases a lot. Mix categories and test your own order, and re-test after model changes.
Should examples contain real customer messages?
Real messages show real phrasing, but remove names, emails and order numbers first. See the data privacy guide.
Few-shot or fine-tuning for a small business chatbot?
Start with examples; they need no dataset and can be edited in minutes. Consider fine-tuning only when you have a large set of correct examples and a measured failure that prompts cannot fix.
Do I need examples if my model is very capable?
Often fewer. A clear instruction may be enough; run Step 5 with and without them and keep only what the test supports.
Related guides
- Prompt engineering for chatbots — the instruction layer that comes before examples.
- Zero-shot vs training data — the ladder for intent classification.
- Chatbot fallback copywriting — writing the refusal example well.
- Chatbot guardrails guide — rules that examples must not contradict.
- Few-shot prompting and prompt engineering — the glossary entries behind the terms used here.
Sources
- Anthropic, Prompting best practices — platform.claude.com/docs/en/build-with-claude/prompt-engineering/multishot-prompting, read 4 October 2026: 3 to 5 examples; relevant, diverse, structured; example tags.
- OpenAI, Prompt engineering — developers.openai.com/api/docs/guides/prompt-engineering, read 4 October 2026: diverse inputs; examples in the developer message; few-shot as an alternative to fine-tuning.
- Google, Prompt design strategies, Gemini API — ai.google.dev/gemini-api/docs/prompting-strategies, read 4 October 2026: always include examples; overfitting warning; identical formatting.
- Zhao et al., Calibrate Before Use, 2021 — arxiv.org/abs/2102.09690, abstract read 4 October 2026: order and example sensitivity.
- Liu et al., What Makes Good In-Context Examples for GPT-3?, 2021 — arxiv.org/abs/2101.06804, abstract read 4 October 2026: similarity-based selection.
- Chatbotscape evaluation methodology. /methodology (continuously updated).
About this guide
Chatbotscape launched in 2026 as an independent review site for chatbot platforms. This guide is editorial guidance built from vendor documentation and research abstracts; the five-step method and the A-B-C test are our own practice, not a published standard. The method does not depend on any platform. To flag an error, write to editorial@chatbotscape.com.
Last updated
5 October 2026 — first published.