LLM Token· Language model basics
LLM Token — The Unit a Language Model Counts, and Bills, in
Quick answer: A token is a chunk of text, usually shorter than a long word and longer than a single letter. A large language model turns your message into tokens, works on those, and writes its reply as tokens. For English, OpenAI's help center gives these rules of thumb: "1 token is approximately 4 characters," "1 token is approximately three-quarters of a word," and "100 tokens are approximately 75 words." They are estimates, not conversions: the real count depends on the model's tokenizer and on the language. For a chatbot, tokens matter for two reasons. They set the context window, the amount the model can consider at once, and they set the bill.
How text becomes tokens
A tokenizer is a fixed table and a set of rules that cuts text into pieces and maps each piece to a number. Most current tokenizers use a method called byte pair encoding, or BPE. The tiktoken README, OpenAI's open-source tokenizer, describes BPE as reversible and lossless, able to handle arbitrary text, and compressing text so that "each token corresponds to about 4 bytes" on average. It also notes that BPE "tends to split words into common subwords," and gives the example of a word like "encoding" being cut into "encod" and "ing".
That has practical consequences you can predict:
- Common short words are usually one token, including the space before them.
- Rare words, brand names, product codes and long numbers are split into several tokens.
- Punctuation, line breaks and symbols can each cost a token.
The model works on the numbers, not on the characters. That is why a model that reads fluently can still stumble on questions like how many letters are in a word: the letters are not what it sees.
Rules of thumb, and where they break
| Source | Rule of thumb for English |
|---|---|
| OpenAI Help Center | 1 token is about 4 characters; about three-quarters of a word; 100 tokens is about 75 words |
| Google Gemini API docs | A token is about 4 characters; 100 tokens is about 60 to 80 English words |
| tiktoken README | About 4 bytes per token on average |
The vendors agree closely for plain English prose, and the sources above all use the word "approximately" or "about". The rule breaks in three places.
Other languages. OpenAI's article warns that "other languages can have different relationships between characters, words, and tokens." A tokenizer trained mostly on English text has fewer whole-word entries for other languages and scripts, so the same meaning often takes more tokens. We do not give a multiplier here because it varies by tokenizer and language; measure a sample of your own text, which the guide shows how to do.
Code, URLs, JSON and numbers. These split into many small pieces, so a tool definition or a pasted spreadsheet row costs more tokens than the same number of English words would.
Different models. Each model family has its own tokenizer, and the count for identical text can differ. Anthropic's token counting documentation says that Claude 4.7 and later models "use a newer tokenizer" and that the same input text "produces approximately 30 percent more tokens than on earlier models." Its advice is to recount prompts against the model you plan to use rather than reuse older counts. The practical lesson: a token count belongs to a model, not to a piece of text.
How to count tokens exactly
Do not estimate when you can count. Each vendor offers a way, and which one to use depends on your model.
- OpenAI. For plain text, use the Tokenizer web tool, or for code use tiktoken with the encoding for your model:
tiktoken.encoding_for_model("gpt-4o")returns the matching tokenizer. For a full request, OpenAI provides an input-token counting endpoint that accepts the same input as the Responses API; its documentation says it "returns the exact count the model will receive," including tokens for message roles and boundaries and for tool definitions. OpenAI's help center warns that a plain-text count may miss message structure, tools, schemas, images and files. - Anthropic. The token counting endpoint "is free to use but subject to requests per minute rate limits" and covers system prompts, tools, images and PDFs. The documentation calls the result an estimate: "In some cases, the actual number of input tokens used when creating a message might differ by a small amount."
- Google. Gemini's
count_tokensmethod returns the input total before a call. After a call, the response'susagefield reports input, output, thinking, cached, tool-use and total token counts.
After any real call, the usage numbers the API returns are the figures you are billed on, so a production bot should log them. That is the cheapest way to see what your traffic really costs, and it is the first measurement in LLM observability. If you want a quick look at how a conversation divides between standing instructions, retrieved passages and history, our LLM token counter breaks it down.
Why tokens set the cost and the limit
The limit. Google's documentation states that "the context window defines the combined limit of input and output tokens." Whatever the model is shown, plus whatever it writes back, has to fit. A long system prompt, many retrieved passages or a long chat history leaves less room for the answer.
The bill. Vendors price tokens separately for input and output, usually per million, and the two rates differ. Some also price cached input differently; see prompt caching. The detail that surprises most teams is that input is not just the customer's last message. In a chat, the platform typically sends the system prompt, any retrieved knowledge-base passages and the earlier turns again with every new message. A short conversation on screen can therefore be many times larger in billed tokens than the visible text. The companion guide works through an invented example in which a six-message chat of 1,140 visible tokens bills about 12,390.
The speed. Replies are generated token by token, so a longer reply takes longer to finish. Length limits such as a maximum-output setting are also in tokens.
Tokens are not the only thing counted
Several other units get confused with tokens:
- Characters or words, which marketing pages use because they are familiar, and which only approximate tokens.
- Messages or conversations, which platforms use for flat plans. Our cost per conversation entry covers turning token prices into that figure.
- Embedding tokens, billed when text is converted to a vector for search. They use their own model and price; see vector embeddings.
- Image, audio and PDF inputs, which are converted to tokens by rules that depend on size, detail or duration. OpenAI's counting guide says image counts reflect size and detail level.
- Thinking tokens, written by reasoning models before they answer. Google lists them as their own line in usage, and Anthropic notes that current-turn thinking counts toward input tokens in a count request.
FAQ
What is a token in AI?
The unit of text a language model processes: a word, part of a word, or a symbol. The model reads and writes sequences of tokens, and its limits and prices are expressed in them.
How many words is 1,000 tokens?
For English prose, roughly 750 words by OpenAI's rule of thumb (100 tokens is about 75 words), and about 600 to 800 by Google's (100 tokens is about 60 to 80 words). Treat it as an estimate and count your own text.
Is a token the same as a word?
No. A short common word is often one token, but longer or rarer words split into pieces, and punctuation and spaces can count too.
Why does the same text have a different token count on different models?
Each model family uses its own tokenizer. Anthropic's documentation, for example, says newer Claude models produce roughly 30 percent more tokens for the same text than earlier ones.
Do other languages use more tokens?
Often, because tokenizers have fewer whole-word entries for some languages and scripts. OpenAI warns that the relationship between characters, words and tokens differs by language. Measure a sample of your actual messages.
Are input and output tokens priced the same?
Not usually. Vendors publish separate rates, so check the current price page for your model rather than relying on a remembered figure.
Related terms
- Large language model — the model that reads and writes the tokens.
- Context window — the token limit on what the model can consider at once.
- Prompt caching — how repeated input tokens can be billed at a lower rate.
- System prompt — the standing instructions that are re-sent, and re-billed, on each message.
- Cost per conversation — the business figure that token prices roll up into.
- Vector embeddings — a separate token-billed step used in retrieval.
Sources
- OpenAI Help Center, What are tokens and how to count them — help.openai.com/en/articles/4936856-what-are-tokens-and-how-to-count-them, read 7 October 2026. Source of the 4-character, three-quarters-of-a-word and 100-tokens-to-75-words rules, the language caveat and the plain-text counting warning.
- OpenAI, Token counting — developers.openai.com/api/docs/guides/token-counting, read 7 October 2026. Source of the input-token counting endpoint and what it includes.
- OpenAI, tiktoken — github.com/openai/tiktoken, read 7 October 2026. Source of the BPE description, the 4-bytes average and encoding_for_model.
- Anthropic, Token counting — platform.claude.com/docs/en/build-with-claude/token-counting, read 7 October 2026. Source of the free-but-rate-limited endpoint, the estimate caveat and the newer-tokenizer note.
- Google, Understand and count tokens, Gemini API — ai.google.dev/gemini-api/docs/tokens, read 7 October 2026. Source of the 4-character and 60-to-80-word rules, count_tokens, the usage field and the combined input-and-output window.
- Chatbotscape evaluation methodology. /methodology (continuously updated).