Chain-of-Thought Prompting· Prompting technique
Chain-of-Thought Prompting — Asking the Model to Show Its Working
Quick answer: Chain-of-thought prompting makes a large language model produce a sequence of reasoning steps before its answer. The 2022 research that named it found large gains on arithmetic, commonsense and symbolic tasks. Later work found the benefit is concentrated in math and logic. Today most major vendors build this behavior into "thinking" or "reasoning" modes, and their guides say that on those models you should rely on the built-in mode rather than add "think step by step" yourself. The technique still matters for models run with thinking off, and for knowing what you are paying for when thinking is on.
A worked example
Here are two prompts for the same invented support question. The numbers are made up for illustration.
Without reasoning:
A customer on the Pro plan ($40 a month) upgraded to Business ($90 a month)
18 days into a 30-day month. What do we charge today? Answer with one number.
With a reasoning instruction:
A customer on the Pro plan ($40 a month) upgraded to Business ($90 a month)
18 days into a 30-day month. Work through it step by step: unused Pro credit,
remaining Business days, then the amount due. Finish with "Amount due: $X".
The second prompt asks for the intermediate quantities the answer depends on. The model writes the credit and the prorated charge before the total, so the final figure is computed from text it has already produced instead of guessed in one jump. That is the whole mechanism, and it is why the technique is associated with arithmetic.
Zero-shot, few-shot and self-consistency versions
| Version | What you supply | Source |
|---|---|---|
| Few-shot CoT | Worked examples that include the steps | Wei et al., 2022 |
| Zero-shot CoT | A phrase such as "Let's think step by step" | Kojima et al., 2022 |
| Self-consistency | Several sampled reasoning paths, then the most common answer | Wang et al., 2022 |
Wei and colleagues showed that "generating a chain of thought — a series of intermediate reasoning steps" improved reasoning in large models, with a 540-billion-parameter model reaching state-of-the-art accuracy on the GSM8K math benchmark from only eight examples. This is few-shot prompting where the examples contain reasoning. Kojima and colleagues then showed that no examples were needed: adding "Let's think step by step" raised InstructGPT accuracy on MultiArith from 17.7% to 78.7% and on GSM8K from 10.4% to 40.7%, per their abstract. Wang and colleagues added self-consistency, which samples diverse reasoning paths and keeps the most consistent answer, reporting for instance +17.9% on GSM8K over greedy decoding. Those figures are from 2022-era models, so treat them as evidence the idea works, not as expected gains today.
What it helps with, and what it does not
Sprague and colleagues (2024) analyzed over 100 papers and ran 20 datasets across 14 models. They report that CoT gives strong gains "primarily on tasks involving math or logic, with much smaller gains on other types of tasks." On MMLU, they found that answering directly was almost identical in accuracy to CoT "unless the question or model's response contains an equals sign." Their conclusion was that CoT can be applied selectively to save inference cost.
For a support bot that means reasoning is worth its tokens on calculations, eligibility rules with several conditions and date arithmetic, and is mostly wasted on "what are your opening hours." Our temperature entry covers the other generation setting teams tune for consistency.
Chain of thought and built-in thinking
Reasoning models and thinking modes generate a hidden or summarized reasoning pass by default, so the prompt-level trick is largely redundant on them.
- OpenAI says of its reasoning models: "Avoid chain-of-thought prompts," because "these models perform reasoning internally, prompting them to 'think step by step' or 'explain your reasoning' is unnecessary."
- Anthropic's prompting guide treats manual CoT as a fallback for when thinking is off, and for its newest models says to rely on thinking instead, at a lower effort level if cost or latency matters.
- Google's Gemini documentation says Gemini models engage in dynamic thinking by default, adjusting effort to the request, and suggests minimal or low thinking for fact retrieval or classification.
The practical rule that follows: match the technique to the mode. Thinking off, a step-by-step instruction may help on a reasoning-heavy task. Thinking on, set the effort level instead of prompting for steps.
Limits in a customer-facing bot
Cost and latency. Reasoning is extra output. Google's documentation says response pricing is the sum of output tokens and thinking tokens, and Anthropic's reports how many of the billed output tokens were internal reasoning. A chatbot that reasons on every message pays that on every message and answers more slowly. See cost per conversation.
Unfaithful explanations. Turpin and colleagues (2023) found that CoT explanations "can be heavily influenced by adding biasing features to model inputs" without the model mentioning them; in their tests, accuracy dropped by as much as 36% on a set of 13 tasks when the inputs were biased toward a wrong answer. A fluent chain of steps is not proof the answer is right or that the steps are the real cause of it.
Not a fix for missing facts. A model can reason carefully from a wrong premise. Reasoning does not remove hallucination; grounding and testing do.
Showing the steps. Raw reasoning text is written for the model, not the customer, and can contain false starts or phrasing you did not approve. We recommend showing customers the answer and, where useful, a short reason you wrote or reviewed, and keeping raw reasoning in logs. That is our editorial judgment, not a vendor requirement.
FAQ
What is chain-of-thought prompting?
A way of getting a language model to write out intermediate reasoning steps before its final answer, using worked examples that include steps or an instruction such as "think step by step."
Does "think step by step" still work?
On models run without thinking it can help on math and logic tasks. On reasoning models OpenAI says it is unnecessary, and Anthropic says to rely on thinking instead for its newest models.
Is chain-of-thought prompting the same as a reasoning model?
No. CoT prompting is something you write in the prompt. A reasoning model or thinking mode generates its own reasoning pass by design, and bills for it.
Does chain of thought make answers more accurate?
Research shows clear gains on math and symbolic tasks and much smaller gains elsewhere. It does not guarantee correctness, and explanations can be unfaithful.
Should I show the reasoning to chatbot users?
Usually not raw. Show the answer, and a reviewed explanation if the customer needs one.
Related terms
- Prompt engineering — the wider craft that CoT belongs to.
- Few-shot prompting — the example-based method that few-shot CoT builds on.
- Zero-shot learning — doing a task from the instruction alone.
- LLM temperature — the sampling setting that interacts with self-consistency.
- AI hallucination — the failure that reasoning does not remove.
Sources
- Anthropic, Prompting best practices — platform.claude.com/docs/en/build-with-claude/prompt-engineering/chain-of-thought, read 5 October 2026. Source of the manual-CoT-as-fallback and rely-on-thinking guidance.
- Anthropic, Extended thinking — platform.claude.com/docs/en/build-with-claude/extended-thinking, read 5 October 2026. Source of the billed-thinking-tokens statement.
- OpenAI, Reasoning best practices — developers.openai.com/api/docs/guides/reasoning-best-practices, read 5 October 2026. Source of the avoid-CoT-prompts guidance.
- Google, Thinking, Gemini API — ai.google.dev/gemini-api/docs/thinking, read 5 October 2026. Source of dynamic thinking, thinking levels and thinking-token pricing.
- Wei et al., Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, 2022 — arxiv.org/abs/2201.11903, abstract read 5 October 2026.
- Kojima et al., Large Language Models are Zero-Shot Reasoners, 2022 — arxiv.org/abs/2205.11916, abstract read 5 October 2026.
- Wang et al., Self-Consistency Improves Chain of Thought Reasoning in Language Models, 2022 — arxiv.org/abs/2203.11171, abstract read 5 October 2026.
- Sprague et al., To CoT or not to CoT?, 2024 — arxiv.org/abs/2409.12183, abstract read 5 October 2026.
- Turpin et al., Language Models Don't Always Say What They Think, 2023 — arxiv.org/abs/2305.04388, abstract read 5 October 2026.
- Chatbotscape evaluation methodology. /methodology (continuously updated).