
How to Estimate Your Chatbot's Token Usage and Monthly AI Bill Before You Launch
Quick answer: Do not estimate from the words customers type. Estimate from what the model is sent. Every message re-sends the system prompt, any retrieved knowledge-base passages and the earlier turns of the chat, so input tokens grow with each turn and usually dwarf the reply. The method: measure your real text with the vendor's counter, build one typical conversation turn by turn, multiply by monthly volume, add a margin, and replace the estimate with your own usage logs in the first two weeks after launch. In the invented example below, a six-message chat of 1,140 visible tokens bills about 12,390, roughly eleven times more.
Step 1: Measure your own text, not an average
Rules of thumb are fine for a sanity check. OpenAI's help center gives about four characters, or three-quarters of a word, per token for English, and Google's documentation says 100 tokens is about 60 to 80 English words. But OpenAI also warns that "other languages can have different relationships between characters, words, and tokens," and Anthropic's documentation says its 4.7 and later models produce "approximately 30 percent more tokens" for the same text than earlier ones. Your text, your language and your model decide the real figure.
So count a sample. Collect 30 real customer messages (anonymized; see the data privacy guide), 30 replies the bot has given or would give, and your actual system prompt. Run each through the counter for the model you will use:
- OpenAI: the Tokenizer web tool, tiktoken, or the input-token counting endpoint, which per OpenAI's guide "returns the exact count the model will receive."
- Anthropic: the token counting endpoint, which the documentation says is "free to use" within rate limits. It describes the result as an estimate that may differ "by a small amount."
- Google: the
count_tokensmethod, which returns the input total before you send.
Write down four averages: tokens per customer message (call it u), tokens per bot reply (a), tokens in the system prompt and tool definitions (S), and tokens in the retrieved passages per turn (R). If your bot is multilingual, take each average per language. Our LLM token counter shows the split between these parts on a pasted conversation.
Step 2: Count what is re-sent every turn
Chat models have no memory between requests. The platform rebuilds the whole prompt each time, which is why the context window fills up. Google's documentation puts it plainly: "the context window defines the combined limit of input and output tokens." For turn n of a conversation, the input is:
input(n) = S + R + (n − 1) × (u + a) + u
That is the standing instructions, the passages fetched for this turn, every earlier customer message and bot reply, and the new message. The output is just a. Two details change the formula:
- If your platform retrieves passages once and keeps them, count R once instead of per turn. If it retrieves again each turn, count it each time. Check which; it changes the total a lot.
- If your platform trims or summarizes old turns, replace the history term with the trimmed size. See manage chatbot context window.
Also remember what else is billed as input: tool and function definitions, which OpenAI's guide says "add tokens to the context," and images or PDFs, which are converted to tokens by size. Reasoning models also produce thinking tokens that appear in usage, so a model that thinks will cost more per turn than the visible reply suggests.
Step 3: Build one typical conversation (worked example)
Every number below is an invented placeholder for illustration, not a measurement or a vendor price. Suppose a support bot averages six customer messages per conversation, and Step 1 gave these averages:
| Part | Tokens |
|---|---|
| System prompt and tools (S) | 800 |
| Retrieved passages per turn (R) | 600 |
| Customer message (u) | 40 |
| Bot reply (a) | 150 |
Using the formula, the input grows each turn:
| Turn | Input tokens | Output tokens |
|---|---|---|
| 1 | 1,440 | 150 |
| 2 | 1,630 | 150 |
| 3 | 1,820 | 150 |
| 4 | 2,010 | 150 |
| 5 | 2,200 | 150 |
| 6 | 2,390 | 150 |
| Total | 11,490 | 900 |
The visible text of the chat is only 6 × (40 + 150) = 1,140 tokens. The billed total is 11,490 + 900 = 12,390, about 10.9 times larger. Most of the gap is the 800-token prompt and 600-token passages being sent six times, plus the history replayed in later turns. This is the number people miss when they estimate from word counts, and it is why a flat-rate plan can look cheap until traffic arrives.
Step 4: Multiply by volume, price and a margin
Take the vendor's current input and output rates for your model, usually quoted per million tokens, and apply:
cost per conversation = (input tokens ÷ 1,000,000 × input rate) + (output tokens ÷ 1,000,000 × output rate)
To show the arithmetic only, assume placeholder rates of $1 per million input tokens and $5 per million output tokens. These are not real prices. Then:
- Input: 11,490 ÷ 1,000,000 × $1 = $0.01149
- Output: 900 ÷ 1,000,000 × $5 = $0.0045
- Per conversation: $0.01599, about 1.6 cents
- At 1,000 conversations a month: about $15.99
Notice that output is under 8 percent of the tokens but about 28 percent of this placeholder bill, because output is typically priced higher. Both columns matter.
Then add a margin. Our own rule, not a standard: add 20 percent for the things a sample misses, such as longer-than-average chats, retries, the occasional pasted document and model changes, and run a second scenario where conversations are twice as long, since history makes cost grow faster than length. If you use prompt caching, apply the cached rate only to the repeated prefix and confirm your prompt meets the model's minimum cacheable length. For the business view, convert the result into cost per conversation and set it beside the value of a resolved chat, as in chatbot ROI quick math.
Step 5: Replace the estimate with real usage logs
An estimate is a plan; usage logs are the truth. API responses report the tokens actually used (Google's usage field, for example, lists input, output, thinking, cached, tool-use and total counts). Log them per conversation from day one, keep no customer text in the log, and in the first two weeks compare:
- Average input and output tokens per conversation against your Step 3 table.
- The 90th-percentile conversation against your doubled-length scenario.
- Tokens per turn by language, if you serve more than one.
If real input is more than 20 percent above your estimate, find out which of S, R or the history term is bigger than assumed; it is almost always the retrieved passages or a system prompt that grew during testing. The monitor LLM chatbot in production guide shows what to record and how to review it weekly.
Common mistakes
Estimating from words typed. Customer messages are the smallest part of the bill.
Using last year's count. A new model may use a different tokenizer, so recount when you change models.
Counting text only. Tool definitions, images, files and thinking tokens all add to the input or output.
Ignoring output price. Output is a minority of tokens but often a large share of cost.
One average for every language. Measure each language you serve.
Never checking the logs. Estimates drift as prompts, knowledge bases and traffic change.
FAQ
How many tokens does a chatbot conversation use?
It depends on your prompt, retrieval and chat length. In our invented example a six-message chat used 12,390 billed tokens, but yours will differ. Build the table in Step 3 from your own averages and verify with logs.
How do I count tokens for my chatbot before launch?
Use the counter for your model: tiktoken or the input-token counting endpoint for OpenAI, the token counting endpoint for Anthropic, or count_tokens for Gemini. Count the system prompt, tools and a sample of real messages.
Why is my chatbot bill higher than I expected?
Usually because input is re-sent every turn: the system prompt, retrieved passages and chat history. Check the usage logs for which part is largest.
Are tokens the same across OpenAI, Anthropic and Google?
No. Each uses its own tokenizer, so identical text can have different counts, and a newer model from the same vendor can change it again.
Do I need to count tokens if I use a no-code platform with a flat price?
Often yes, at least roughly. Flat plans usually include a fair-use allowance or charge for extra usage, and a bring-your-own-key setup bills you directly by tokens. See the pricing models guide.
Is the vendor's token count exact?
OpenAI describes its input-token endpoint as returning the exact count; Anthropic describes its count as an estimate that can differ slightly. Whichever you use, trust the usage figures returned by real calls over any pre-call count.
Related guides
- Reduce chatbot costs guide — the levers to pull once you know where the tokens go.
- Manage the chatbot context window — trimming and summarizing history.
- Monitor an LLM chatbot in production — logging usage and reviewing it weekly.
- Voice AI cost guide — the same estimating habit for voice, where minutes matter too.
- LLM token and prompt caching — the glossary entries behind the terms used here.
Sources
- OpenAI Help Center, What are tokens and how to count them — help.openai.com/en/articles/4936856-what-are-tokens-and-how-to-count-them, read 7 October 2026: token-to-character and word rules of thumb; language caveat.
- OpenAI, Token counting — developers.openai.com/api/docs/guides/token-counting, read 7 October 2026: input-token counting endpoint; tools add tokens.
- Anthropic, Token counting — platform.claude.com/docs/en/build-with-claude/token-counting, read 7 October 2026: free endpoint, estimate caveat, newer tokenizer note.
- Google, Understand and count tokens, Gemini API — ai.google.dev/gemini-api/docs/tokens, read 7 October 2026: count_tokens, usage field, combined window, words-per-token range.
- Chatbotscape evaluation methodology. /methodology (continuously updated).
About this guide
Chatbotscape launched in 2026 as an independent review site for chatbot platforms. This guide is editorial guidance built from vendor documentation; the five-step method, the 20 percent margin and the worked example are our own practice with invented numbers, not a published standard or a price quote. The method does not depend on any platform. To flag an error, write to editorial@chatbotscape.com.
Last updated
8 October 2026 — first published.