Skip to content
Chatbotscape
Anthropic's, OpenAI's and Google's documentation and the original research papers were read on 5 October 2026. Vendor advice on reasoning changes with model releases, so check the current page for the model your bot runs on.
Chain-of-Thought Prompting· Prompting technique
Chain-of-thought (CoT) prompting means getting a language model to write out intermediate reasoning steps before it gives a final answer, either by showing worked examples that include the steps or by instructing it to reason step by step. The steps are generated text, not a view into the model's internals.
By Chatbotscape Editorial· Methodology· Published 6 October 2026· Updated 6 October 2026

Chain-of-Thought Prompting — Asking the Model to Show Its Working

Quick answer: Chain-of-thought prompting makes a large language model produce a sequence of reasoning steps before its answer. The 2022 research that named it found large gains on arithmetic, commonsense and symbolic tasks. Later work found the benefit is concentrated in math and logic. Today most major vendors build this behavior into "thinking" or "reasoning" modes, and their guides say that on those models you should rely on the built-in mode rather than add "think step by step" yourself. The technique still matters for models run with thinking off, and for knowing what you are paying for when thinking is on.

A worked example

Here are two prompts for the same invented support question. The numbers are made up for illustration.

Without reasoning:
A customer on the Pro plan ($40 a month) upgraded to Business ($90 a month)
18 days into a 30-day month. What do we charge today? Answer with one number.

With a reasoning instruction:
A customer on the Pro plan ($40 a month) upgraded to Business ($90 a month)
18 days into a 30-day month. Work through it step by step: unused Pro credit,
remaining Business days, then the amount due. Finish with "Amount due: $X".

The second prompt asks for the intermediate quantities the answer depends on. The model writes the credit and the prorated charge before the total, so the final figure is computed from text it has already produced instead of guessed in one jump. That is the whole mechanism, and it is why the technique is associated with arithmetic.

Zero-shot, few-shot and self-consistency versions

VersionWhat you supplySource
Few-shot CoTWorked examples that include the stepsWei et al., 2022
Zero-shot CoTA phrase such as "Let's think step by step"Kojima et al., 2022
Self-consistencySeveral sampled reasoning paths, then the most common answerWang et al., 2022

Wei and colleagues showed that "generating a chain of thought — a series of intermediate reasoning steps" improved reasoning in large models, with a 540-billion-parameter model reaching state-of-the-art accuracy on the GSM8K math benchmark from only eight examples. This is few-shot prompting where the examples contain reasoning. Kojima and colleagues then showed that no examples were needed: adding "Let's think step by step" raised InstructGPT accuracy on MultiArith from 17.7% to 78.7% and on GSM8K from 10.4% to 40.7%, per their abstract. Wang and colleagues added self-consistency, which samples diverse reasoning paths and keeps the most consistent answer, reporting for instance +17.9% on GSM8K over greedy decoding. Those figures are from 2022-era models, so treat them as evidence the idea works, not as expected gains today.

What it helps with, and what it does not

Sprague and colleagues (2024) analyzed over 100 papers and ran 20 datasets across 14 models. They report that CoT gives strong gains "primarily on tasks involving math or logic, with much smaller gains on other types of tasks." On MMLU, they found that answering directly was almost identical in accuracy to CoT "unless the question or model's response contains an equals sign." Their conclusion was that CoT can be applied selectively to save inference cost.

For a support bot that means reasoning is worth its tokens on calculations, eligibility rules with several conditions and date arithmetic, and is mostly wasted on "what are your opening hours." Our temperature entry covers the other generation setting teams tune for consistency.

Chain of thought and built-in thinking

Reasoning models and thinking modes generate a hidden or summarized reasoning pass by default, so the prompt-level trick is largely redundant on them.

  • OpenAI says of its reasoning models: "Avoid chain-of-thought prompts," because "these models perform reasoning internally, prompting them to 'think step by step' or 'explain your reasoning' is unnecessary."
  • Anthropic's prompting guide treats manual CoT as a fallback for when thinking is off, and for its newest models says to rely on thinking instead, at a lower effort level if cost or latency matters.
  • Google's Gemini documentation says Gemini models engage in dynamic thinking by default, adjusting effort to the request, and suggests minimal or low thinking for fact retrieval or classification.

The practical rule that follows: match the technique to the mode. Thinking off, a step-by-step instruction may help on a reasoning-heavy task. Thinking on, set the effort level instead of prompting for steps.

Limits in a customer-facing bot

Cost and latency. Reasoning is extra output. Google's documentation says response pricing is the sum of output tokens and thinking tokens, and Anthropic's reports how many of the billed output tokens were internal reasoning. A chatbot that reasons on every message pays that on every message and answers more slowly. See cost per conversation.

Unfaithful explanations. Turpin and colleagues (2023) found that CoT explanations "can be heavily influenced by adding biasing features to model inputs" without the model mentioning them; in their tests, accuracy dropped by as much as 36% on a set of 13 tasks when the inputs were biased toward a wrong answer. A fluent chain of steps is not proof the answer is right or that the steps are the real cause of it.

Not a fix for missing facts. A model can reason carefully from a wrong premise. Reasoning does not remove hallucination; grounding and testing do.

Showing the steps. Raw reasoning text is written for the model, not the customer, and can contain false starts or phrasing you did not approve. We recommend showing customers the answer and, where useful, a short reason you wrote or reviewed, and keeping raw reasoning in logs. That is our editorial judgment, not a vendor requirement.

FAQ

What is chain-of-thought prompting?

A way of getting a language model to write out intermediate reasoning steps before its final answer, using worked examples that include steps or an instruction such as "think step by step."

Does "think step by step" still work?

On models run without thinking it can help on math and logic tasks. On reasoning models OpenAI says it is unnecessary, and Anthropic says to rely on thinking instead for its newest models.

Is chain-of-thought prompting the same as a reasoning model?

No. CoT prompting is something you write in the prompt. A reasoning model or thinking mode generates its own reasoning pass by design, and bills for it.

Does chain of thought make answers more accurate?

Research shows clear gains on math and symbolic tasks and much smaller gains elsewhere. It does not guarantee correctness, and explanations can be unfaithful.

Should I show the reasoning to chatbot users?

Usually not raw. Show the answer, and a reviewed explanation if the customer needs one.

Sources