
Step-by-Step Reasoning for a Support Chatbot
When It Helps, How to Prompt It, and How to Test the Cost
Quick answer: First find out whether your model already reasons. If it has a built-in thinking mode, OpenAI says extra "think step by step" prompting is unnecessary and Anthropic says to rely on thinking instead, so set the effort level rather than writing steps. If thinking is off, add a short reasoning instruction only to the questions that involve calculation or several conditions, because research finds the benefit is concentrated in math and logic. Keep the raw reasoning out of the customer's chat window, and prove the change helped by running the same 30 real messages with and without it and comparing accuracy, cost and response time.
Step 1: Find out whether your model already reasons
Check the model setting in your platform or API call. Three situations cover most bots.
| Your setup | What the vendor docs say | What to do |
|---|---|---|
| A reasoning or thinking mode is on | OpenAI: avoid chain-of-thought prompts, the models reason internally. Google: Gemini engages in dynamic thinking by default | Do not add step-by-step wording. Choose the effort or thinking level |
| A thinking-capable model with thinking switched off | Anthropic: manual chain of thought is a fallback when thinking is off | Add a reasoning instruction only where Step 2 says it pays |
| A fast non-reasoning model | Research on CoT (Wei 2022, Kojima 2022) was done on this kind of model | Same as above, and test it in Step 5 |
Many no-code platforms hide this choice. If yours does not show which model or thinking level it uses, you cannot tell whether your added instruction is redundant, so treat Step 5 as mandatory. Note also that Anthropic's guide advises a lower effort level when cost or latency matters, and Google suggests minimal or low thinking for fact retrieval or classification.
Step 2: List the questions that actually need reasoning
Sprague and colleagues (2024) found CoT helps "primarily on tasks involving math or logic." Sort your 30 most common customer messages into three bins.
- Reason: proration and refund math, shipping-date arithmetic, eligibility with two or more conditions ("returns allowed if under 30 days and unopened, unless sale item").
- Look up: opening hours, address, plan features, order status. These need a correct source, not steps; see retrieval-augmented generation.
- Route: complaints, legal or safety topics, anything the bot should hand off. Reasoning does not make these safe to answer.
If the "reason" bin is small, say five messages in thirty, you may not need chain-of-thought at all. A cheaper fix is to move the calculation out of the model, for example with function calling to a pricing function, and let the model only explain the result. Sprague and colleagues note CoT "underperforms relative to using a symbolic solver."
Step 3: Write the instruction
For each "reason" question type, give the model the quantities to work out in order, and a fixed last line. An invented example for a billing bot:
When a customer asks what they owe after changing plans mid-month:
1. State the old plan price, the new plan price and the days used.
2. Work out the unused credit and the prorated charge for the new plan.
3. Subtract the credit from the charge.
End the reasoning with a line "Amount due: $X" and nothing after it.
Three habits make this reliable. Name the intermediate quantities instead of writing only "think carefully," so each step is checkable. Fix the final line so your system can read the answer. And if you want the model to check itself, Anthropic's guide suggests appending something like "Before you finish, verify your answer," which costs little. If you are also using examples, put them in the same format, as in our few-shot examples guide; examples that include the steps are the original few-shot form of chain of thought.
Keep the instruction in the system prompt or the relevant flow node rather than appending it to every message, so the "look up" questions stay short and cheap. For the wider prompt structure see prompt engineering for chatbots.
Step 4: Keep raw reasoning off the customer's screen
A model's written steps are working notes. They can contain a wrong first attempt, internal wording or a step that contradicts policy. Turpin and colleagues (2023) also showed that CoT explanations can be influenced by biasing features without the model mentioning them, so the steps are not a reliable account of why the model answered as it did.
Our recommendation, which is editorial judgment and not a vendor rule:
- Send the customer the final answer and, if it helps, a short explanation written in your own template ("We credited the unused 12 days of your Pro plan").
- Keep the full reasoning in your logs for review and for the test in Step 5.
- If your platform returns reasoning and answer together, separate them with a fixed marker and strip everything before it. Test that the marker never appears in a reply.
Vendor settings differ here: Google documents thought summaries that can be turned on, and Anthropic's extended-thinking documentation refers to summarized thinking. Check what your model returns before you decide what to filter, and do not assume a summary is safe to display verbatim. Regulated industries should also ask counsel before showing or storing model reasoning; we do not give legal advice. Our guardrails guide covers rules the reply must still follow.
Step 5: Test with and without reasoning (our method)
This is Chatbotscape's own procedure, not a published standard.
- Take 30 real customer messages, with the 10 or so from your "reason" bin included. Write the correct answer for each. This is a small golden dataset.
- Run all 30 with your current prompt and record: correct or not, tokens used, and seconds to reply.
- Add the reasoning instruction (or raise the thinking level) and run the same 30 again.
- Compare the "reason" bin and the other 20 separately. Reasoning should improve the first group; if it does not, remove it.
- Decide per group. We suggest keeping reasoning only where it fixes at least two of ten wrong answers and the extra cost and delay are acceptable to you. That threshold is a starting point; set your own.
Use LLM-as-a-judge only to scale this up, and spot-check its grades by hand. If you vary the temperature during the test, change one thing at a time; the temperature settings guide explains why. Repeat the test when you change models, since vendor guidance on reasoning has changed across releases.
Cost is real: Google states that response pricing is the sum of output tokens and thinking tokens, and Anthropic's documentation reports how many billed output tokens were internal reasoning. Multiply the per-message increase by your monthly volume before you roll out, and compare it with the savings guidance in reduce chatbot costs.
Common mistakes
Adding "think step by step" to every prompt. On a model that already thinks it adds nothing the vendors recommend, and on a lookup question it adds cost.
Showing the reasoning to customers. It exposes notes you did not write and invites screenshots of a mid-thought error.
Trusting the steps as proof. A fluent chain can still reach a wrong answer from a wrong premise.
Using reasoning to patch a missing fact. The model will reason confidently around a gap in your knowledge base.
Testing on easy messages only. If your 30 messages contain no calculations, the test will show no benefit and you will conclude the wrong thing.
FAQ
Does chain-of-thought prompting work on chatbots?
On calculation and multi-condition questions, yes, mainly on models that are not already reasoning. On lookups it adds cost without benefit.
Should I put "think step by step" in my system prompt?
Not globally. Add it only to the question types that need it, and skip it if your model has a thinking mode.
What is the difference between chain of thought and a reasoning model?
Chain of thought is wording you add to the prompt. A reasoning model does its own reasoning pass and bills for it.
Will it slow my bot down?
Usually yes, because the model writes more before it answers. Measure response time in Step 5.
Should customers see the reasoning?
We recommend showing the answer and a short reviewed explanation, and keeping raw reasoning in logs.
Related guides
- Prompt engineering for chatbots — the instruction layer this builds on.
- Few-shot examples for chatbots — examples that include the steps.
- Chatbot temperature settings — the other generation setting to test.
- Reduce chatbot hallucinations — when the problem is facts, not reasoning.
- Chain-of-thought prompting and prompt engineering — the glossary entries behind the terms used here.
Sources
- Anthropic, Prompting best practices — platform.claude.com/docs/en/build-with-claude/prompt-engineering/chain-of-thought, read 5 October 2026: manual chain of thought as a fallback; rely on thinking; self-check suggestion; lower effort for cost.
- Anthropic, Extended thinking — platform.claude.com/docs/en/build-with-claude/extended-thinking, read 5 October 2026: thinking tokens count within billed output tokens; summarized thinking.
- OpenAI, Reasoning best practices — developers.openai.com/api/docs/guides/reasoning-best-practices, read 5 October 2026: avoid chain-of-thought prompts on reasoning models.
- Google, Thinking, Gemini API — ai.google.dev/gemini-api/docs/thinking, read 5 October 2026: dynamic thinking, thinking levels, thought summaries, thinking-token pricing.
- Sprague et al., To CoT or not to CoT?, 2024 — arxiv.org/abs/2409.12183, abstract read 5 October 2026.
- Turpin et al., Language Models Don't Always Say What They Think, 2023 — arxiv.org/abs/2305.04388, abstract read 5 October 2026.
- Wei et al., 2022 — arxiv.org/abs/2201.11903 and Kojima et al., 2022 — arxiv.org/abs/2205.11916, abstracts read 5 October 2026.
- Chatbotscape evaluation methodology. /methodology (continuously updated).
About this guide
Chatbotscape launched in 2026 as an independent review site for chatbot platforms. This guide is editorial guidance built from vendor documentation and research abstracts; the five-step method and the with-and-without test are our own practice, not a published standard. The method does not depend on any platform. To flag an error, write to editorial@chatbotscape.com.
Last updated
6 October 2026 — first published.