Few-Shot Prompting· Prompting technique
Few-Shot Prompting — Showing the Model What a Good Answer Looks Like
Quick answer: Few-shot prompting is giving a large language model a few finished examples before the real question. "Shot" means example. Zero-shot has none, one-shot has one, few-shot has several. It works because the model continues whatever pattern it sees, so examples are an efficient way to fix the output format, the tone and the category names. It is a prompt technique, not training: the model's weights do not change, and you pay for the examples again on every call.
A worked example
Here is a few-shot prompt for sorting incoming messages. The messages are invented for illustration.
Classify the customer message as billing, shipping, or other.
Message: "I was charged twice for order 4411"
Category: billing
Message: "Where is my parcel? It said Tuesday"
Category: shipping
Message: "Do you have a store in Leeds?"
Category: other
Message: "My card was declined at checkout but I still got an email"
Category:
The first three message-and-category pairs are the shots. The model is not told what "billing" means in a legal sense; it infers the task, the label set and the layout from the pattern and completes the last line. Remove the three examples and you have the same task as zero-shot, which relies on the instruction alone.
How it differs from nearby techniques
| Technique | What you supply | Persists? | Typical cost |
|---|---|---|---|
| Zero-shot | An instruction only | No | Shortest prompt |
| One-shot | An instruction and one example | No | Slightly longer prompt |
| Few-shot | An instruction and several examples | No | Longer prompt on every call |
| Fine-tuning | A training dataset, used once | Yes, in a new model version | Dataset, training run, upkeep |
The original few-shot result comes from the GPT-3 paper by Brown and colleagues (2020), which applied the model "without any gradient updates or fine-tuning, with tasks and few-shot demonstrations specified purely via text interaction." That sentence is the whole idea: the learning happens inside one request, which is why researchers call it in-context learning. Our zero-shot learning entry covers the other end of the scale.
What the examples actually teach
It is tempting to think each example teaches the model a fact. Research suggests it teaches the shape of the task. Min and colleagues (2022) tested this on classification and multiple-choice tasks across twelve models including GPT-3, and found that "randomly replacing labels in the demonstrations barely hurts performance." What demonstrations supplied, in their analysis, was the label space, the distribution of the input text and the format of the sequence.
Two cautions keep that from being over-read. The finding is about classification and multiple-choice tasks, not about a bot that quotes your refund policy, where a wrong example is simply a wrong answer to copy. And later work found that scale changes the picture: Wei and colleagues (2023) report that "overriding semantic priors is an emergent ability of model scale," meaning larger models do follow examples even when they contradict what the model learned in pre-training. The practical reading is to make every example correct and to treat the format and category names it shows as the part that matters most.
Which examples, and in what order
Few-shot results are sensitive to choices that look cosmetic.
- Order. Zhao and colleagues (2021) found that "the choice of prompt format, training examples, and even the order of the training examples can cause accuracy to vary from near chance to near state-of-the-art," and noted a bias toward answers placed near the end of the prompt. Lu and colleagues (2022) report that order "can make the difference between near state-of-the-art and random guess performance." These results come from older models; current models are likely less fragile, but the safe habit is to test your own order rather than assume.
- Selection. Liu and colleagues (2021) showed that choosing examples by semantic similarity to the incoming message beat choosing them at random, reporting gains of 41.9% on one table-to-text benchmark and 45.5% on one open-domain question-answering benchmark. Retrieving the closest examples per message is the idea behind dynamic example selection, and it uses the same machinery as semantic search.
- Diversity. Anthropic's guide says examples should be relevant, diverse and structured, and recommends "3–5 examples for best results." OpenAI's guide says to "show a diverse range of possible inputs with the desired outputs."
Limits in a customer-facing bot
Cost and length. Every example is sent on every request. A long block of examples raises the per-message cost and uses space in the context window that retrieved documents could use.
Too many can hurt. Google's Gemini guide recommends always including examples, but warns that "if you include too many examples, the model may start to overfit the response to the examples." In practice that looks like a bot repeating an example's wording back to customers.
Format drift. Google's guide also says to "make sure that the structure and formatting of few-shot examples are the same." One example that answers in a different layout teaches the model that layout is optional.
Not a fact source. An example that includes a price teaches the model a price. When the price changes, the bot keeps quoting it. Keep changing facts in the knowledge base, not in the examples.
Instruction versus example. Examples sit beside the instructions in a system prompt and can pull against them. If the rules say "never promise a refund" and one example does, the model sees a conflict, and the example is the more concrete signal.
FAQ
What is few-shot prompting?
Putting a few input-and-output examples in the prompt so the model imitates the pattern on a new input. The model is not retrained and the examples last only for that request.
What is the difference between zero-shot and few-shot prompting?
Zero-shot gives the model an instruction and no examples; few-shot adds several. Few-shot usually helps most when the output needs a specific format or a fine distinction between categories.
How many examples should a few-shot prompt have?
Anthropic recommends 3 to 5; Google says to experiment, and warns that too many can cause overfitting. There is no universal number; test on your own messages.
Is few-shot prompting the same as fine-tuning?
No. Fine-tuning changes the model through a training run and persists. Few-shot changes nothing about the model and is resent each time. OpenAI's guide presents few-shot as an alternative to fine-tuning when steering a model toward a new task.
Does the order of examples matter?
It can. Research on older models found large swings from reordering, with a bias toward answers near the end. Test your own order.
Related terms
- Prompt engineering — the wider craft of which few-shot examples are one tool.
- Zero-shot learning — the same task with no examples in the prompt.
- System prompt — where most platforms let you place instructions and examples.
- Fine-tuning — the heavier option when examples no longer fit in a prompt.
- Golden dataset — the fixed message set that shows whether examples helped.
Sources
- Anthropic, Prompting best practices — platform.claude.com/docs/en/build-with-claude/prompt-engineering/multishot-prompting, read 4 October 2026. Source of "3–5 examples" and the relevant, diverse, structured guidance.
- OpenAI, Prompt engineering — developers.openai.com/api/docs/guides/prompt-engineering, read 4 October 2026. Source of the diverse-inputs advice and the few-shot-instead-of-fine-tuning framing.
- Google, Prompt design strategies, Gemini API — ai.google.dev/gemini-api/docs/prompting-strategies, read 4 October 2026. Source of the overfitting warning and the consistent-formatting advice.
- Brown et al., Language Models are Few-Shot Learners, 2020 — arxiv.org/abs/2005.14165, abstract read 4 October 2026.
- Min et al., Rethinking the Role of Demonstrations, 2022 — arxiv.org/abs/2202.12837, abstract read 4 October 2026.
- Zhao et al., Calibrate Before Use, 2021 — arxiv.org/abs/2102.09690, abstract read 4 October 2026.
- Lu et al., Fantastically Ordered Prompts and Where to Find Them, 2022 — arxiv.org/abs/2104.08786, abstract read 4 October 2026.
- Liu et al., What Makes Good In-Context Examples for GPT-3?, 2021 — arxiv.org/abs/2101.06804, abstract read 4 October 2026.
- Wei et al., Larger language models do in-context learning differently, 2023 — arxiv.org/abs/2303.03846, abstract read 4 October 2026.
- Chatbotscape evaluation methodology. /methodology (continuously updated).