Skip to content
Chatbotscape
OpenAI's, Anthropic's and Google's token documentation and the tiktoken README were read on 7 October 2026. Tokenizers change between model generations, so recount your own text against the model you actually run.
LLM Token· Language model basics
A token is the unit of text a language model reads and writes: a whole short word, a piece of a longer word, a number fragment or a punctuation mark. The model never sees letters or words, only a sequence of token IDs. Limits, speed and price are all measured in tokens.
By Chatbotscape Editorial· Methodology· Published 8 October 2026· Updated 8 October 2026

LLM Token — The Unit a Language Model Counts, and Bills, in

Quick answer: A token is a chunk of text, usually shorter than a long word and longer than a single letter. A large language model turns your message into tokens, works on those, and writes its reply as tokens. For English, OpenAI's help center gives these rules of thumb: "1 token is approximately 4 characters," "1 token is approximately three-quarters of a word," and "100 tokens are approximately 75 words." They are estimates, not conversions: the real count depends on the model's tokenizer and on the language. For a chatbot, tokens matter for two reasons. They set the context window, the amount the model can consider at once, and they set the bill.

How text becomes tokens

A tokenizer is a fixed table and a set of rules that cuts text into pieces and maps each piece to a number. Most current tokenizers use a method called byte pair encoding, or BPE. The tiktoken README, OpenAI's open-source tokenizer, describes BPE as reversible and lossless, able to handle arbitrary text, and compressing text so that "each token corresponds to about 4 bytes" on average. It also notes that BPE "tends to split words into common subwords," and gives the example of a word like "encoding" being cut into "encod" and "ing".

That has practical consequences you can predict:

  • Common short words are usually one token, including the space before them.
  • Rare words, brand names, product codes and long numbers are split into several tokens.
  • Punctuation, line breaks and symbols can each cost a token.

The model works on the numbers, not on the characters. That is why a model that reads fluently can still stumble on questions like how many letters are in a word: the letters are not what it sees.

Rules of thumb, and where they break

SourceRule of thumb for English
OpenAI Help Center1 token is about 4 characters; about three-quarters of a word; 100 tokens is about 75 words
Google Gemini API docsA token is about 4 characters; 100 tokens is about 60 to 80 English words
tiktoken READMEAbout 4 bytes per token on average

The vendors agree closely for plain English prose, and the sources above all use the word "approximately" or "about". The rule breaks in three places.

Other languages. OpenAI's article warns that "other languages can have different relationships between characters, words, and tokens." A tokenizer trained mostly on English text has fewer whole-word entries for other languages and scripts, so the same meaning often takes more tokens. We do not give a multiplier here because it varies by tokenizer and language; measure a sample of your own text, which the guide shows how to do.

Code, URLs, JSON and numbers. These split into many small pieces, so a tool definition or a pasted spreadsheet row costs more tokens than the same number of English words would.

Different models. Each model family has its own tokenizer, and the count for identical text can differ. Anthropic's token counting documentation says that Claude 4.7 and later models "use a newer tokenizer" and that the same input text "produces approximately 30 percent more tokens than on earlier models." Its advice is to recount prompts against the model you plan to use rather than reuse older counts. The practical lesson: a token count belongs to a model, not to a piece of text.

How to count tokens exactly

Do not estimate when you can count. Each vendor offers a way, and which one to use depends on your model.

  • OpenAI. For plain text, use the Tokenizer web tool, or for code use tiktoken with the encoding for your model: tiktoken.encoding_for_model("gpt-4o") returns the matching tokenizer. For a full request, OpenAI provides an input-token counting endpoint that accepts the same input as the Responses API; its documentation says it "returns the exact count the model will receive," including tokens for message roles and boundaries and for tool definitions. OpenAI's help center warns that a plain-text count may miss message structure, tools, schemas, images and files.
  • Anthropic. The token counting endpoint "is free to use but subject to requests per minute rate limits" and covers system prompts, tools, images and PDFs. The documentation calls the result an estimate: "In some cases, the actual number of input tokens used when creating a message might differ by a small amount."
  • Google. Gemini's count_tokens method returns the input total before a call. After a call, the response's usage field reports input, output, thinking, cached, tool-use and total token counts.

After any real call, the usage numbers the API returns are the figures you are billed on, so a production bot should log them. That is the cheapest way to see what your traffic really costs, and it is the first measurement in LLM observability. If you want a quick look at how a conversation divides between standing instructions, retrieved passages and history, our LLM token counter breaks it down.

Why tokens set the cost and the limit

The limit. Google's documentation states that "the context window defines the combined limit of input and output tokens." Whatever the model is shown, plus whatever it writes back, has to fit. A long system prompt, many retrieved passages or a long chat history leaves less room for the answer.

The bill. Vendors price tokens separately for input and output, usually per million, and the two rates differ. Some also price cached input differently; see prompt caching. The detail that surprises most teams is that input is not just the customer's last message. In a chat, the platform typically sends the system prompt, any retrieved knowledge-base passages and the earlier turns again with every new message. A short conversation on screen can therefore be many times larger in billed tokens than the visible text. The companion guide works through an invented example in which a six-message chat of 1,140 visible tokens bills about 12,390.

The speed. Replies are generated token by token, so a longer reply takes longer to finish. Length limits such as a maximum-output setting are also in tokens.

Tokens are not the only thing counted

Several other units get confused with tokens:

  • Characters or words, which marketing pages use because they are familiar, and which only approximate tokens.
  • Messages or conversations, which platforms use for flat plans. Our cost per conversation entry covers turning token prices into that figure.
  • Embedding tokens, billed when text is converted to a vector for search. They use their own model and price; see vector embeddings.
  • Image, audio and PDF inputs, which are converted to tokens by rules that depend on size, detail or duration. OpenAI's counting guide says image counts reflect size and detail level.
  • Thinking tokens, written by reasoning models before they answer. Google lists them as their own line in usage, and Anthropic notes that current-turn thinking counts toward input tokens in a count request.

FAQ

What is a token in AI?

The unit of text a language model processes: a word, part of a word, or a symbol. The model reads and writes sequences of tokens, and its limits and prices are expressed in them.

How many words is 1,000 tokens?

For English prose, roughly 750 words by OpenAI's rule of thumb (100 tokens is about 75 words), and about 600 to 800 by Google's (100 tokens is about 60 to 80 words). Treat it as an estimate and count your own text.

Is a token the same as a word?

No. A short common word is often one token, but longer or rarer words split into pieces, and punctuation and spaces can count too.

Why does the same text have a different token count on different models?

Each model family uses its own tokenizer. Anthropic's documentation, for example, says newer Claude models produce roughly 30 percent more tokens for the same text than earlier ones.

Do other languages use more tokens?

Often, because tokenizers have fewer whole-word entries for some languages and scripts. OpenAI warns that the relationship between characters, words and tokens differs by language. Measure a sample of your actual messages.

Are input and output tokens priced the same?

Not usually. Vendors publish separate rates, so check the current price page for your model rather than relying on a remembered figure.

Sources