The short answer
Quick answer: Language models do not read letters or words. Text is first cut into tokens, pieces that may be a whole word, part of a word, a single character or a punctuation mark, and each token is replaced by an ID number. The model only ever sees and produces those numbers. Most models use a subword method such as byte pair encoding (BPE), which keeps common words whole and splits rare words into smaller known pieces. Tokenization matters in practice because context limits and prices are counted in tokens, and because it explains quirks such as models miscounting letters or fumbling long numbers.
Why not just use words or letters?
A neural network works on numbers, so text must be mapped to a fixed list of units, the vocabulary. There are three obvious choices.
| Unit | Vocabulary size | Problem |
|---|---|---|
| Words | Millions | Cannot handle new words, typos, names or code; a huge table |
| Characters | A few hundred | Sequences become very long, and the model must learn spelling from scratch |
| Subwords | Tens of thousands to a couple of hundred thousand | A workable compromise |
Subword tokenization keeps frequent words as single tokens and builds rare words from pieces:
"cat" -> ["cat"]
"unbelievably" -> ["un", "believ", "ably"]
"Zxqvy" -> ["Z", "x", "q", "v", "y"]
Nothing is ever "unknown": any string can be built from smaller parts, down to single bytes.
How byte pair encoding works
Byte pair encoding began as a data compression technique. As a tokenizer, it is trained once on a large body of text.
Training the vocabulary
- Start with a vocabulary of single characters, or of the 256 possible bytes.
- Count every adjacent pair of tokens in the training text.
- Merge the most frequent pair into a new token and add it to the vocabulary.
- Repeat until the vocabulary reaches the target size.
A toy example with the words low, lower, lowest, newer, newest:
l+oappears often, so merge tolo.lo+wis now frequent, so merge tolow.e+rbecomeser;e+s+tbecomesest.
The result is a vocabulary that includes low, er and est, so lowest is two tokens and a new word like slowest can still be built from known parts.
Using it
To tokenize new text, apply the learned merges in the same order. To turn tokens back into text, just join the pieces. The process is exactly reversible.
OpenAI publishes its tokenizer as the open-source library tiktoken. Other families use related methods, such as WordPiece (BERT) and SentencePiece (many multilingual models). Each model has its own vocabulary; token IDs from one model mean nothing to another.
Details that surprise people
- Spaces belong to tokens. In most BPE vocabularies, the space is attached to the start of the following word.
" hello"and"hello"are different tokens. - Capitalisation matters.
"Hello","hello"and"HELLO"may tokenize differently. - Numbers are chopped arbitrarily.
"1234567"might become["123", "456", "7"], with no respect for place value. - Languages differ. Text in languages with less representation in the tokenizer's training data, or in non-Latin scripts, often needs more tokens for the same meaning. See why Unicode exists for how text is stored as bytes in the first place.
- Code and unusual strings such as random IDs and URLs break into many small tokens.
- Special tokens mark things like the start of a message or the end of the text. They are not ordinary text.
Rules of thumb for counting
For English text with common tokenizers:
- 1 token is roughly 4 characters.
- 1 token is roughly three-quarters of a word.
- 100 tokens is about 75 words.
These are approximations. Always count with the actual tokenizer for the model you are using, especially for code or other languages.
Why tokenization affects behaviour
Limits and cost are measured in tokens
- A model's context window is a number of tokens, covering both the input and the output.
- API pricing is per token, usually with different rates for input and output.
- Output length limits are in tokens.
So verbose formats cost more. The same data as compact text may use far fewer tokens than deeply nested JSON with long key names.
Letter-level tasks are hard
Ask a model how many times a letter appears in a word, and it may get it wrong. The model never saw the letters. It saw one or two token IDs, and has to rely on what it learned indirectly about how those tokens are spelled. Reversing a string or writing text with exact letter constraints is hard for the same reason.
Arithmetic on long numbers is unreliable
Digits are grouped into tokens inconsistently, so the model cannot line up columns the way a person or a calculator does. It also predicts the answer left to right, most significant digit first, while carrying works right to left. Models are better at this than they used to be, particularly when allowed to work step by step, but a calculator tool is the dependable solution.
Small prompt changes shift the output
An extra space, a different capital letter or a typo changes the token sequence, which changes every probability the model computes afterwards. See how LLMs predict the next word.
Rare words are built from parts
Because unknown words are assembled from familiar pieces, models can often handle new terminology sensibly, and can also be misled when the pieces suggest the wrong meaning.
What happens after tokenization
Each token ID is looked up in a table to fetch its embedding, a vector of numbers that the model actually processes. See how embeddings work. The tokenizer is fixed before training; the embeddings are learned during it.
Not only text
The same idea extends to other kinds of data. Images are cut into small patches, each treated as a token. Audio is converted into sequences of discrete codes. That is how a single transformer can handle text, images and sound together.
Practical tips
- Count tokens before sending large prompts, using the provider's tokenizer or token-counting endpoint.
- Trim what the model does not need. Boilerplate, repeated instructions and unused fields all cost tokens.
- Prefer compact formats for large data.
- Do not rely on the model for exact character counts or long arithmetic. Use code or tools.
- Watch for truncation. If input exceeds the context window, something has to be dropped.
Frequently asked questions
What is a token?
A unit of text that a language model processes: a word, part of a word, a character or punctuation, mapped to a number.
How many tokens is a word?
In English, a word averages about 1.3 tokens. Common short words are one token; long or rare words are several.
Why can't an LLM count letters reliably?
It processes tokens, not characters. A word may be a single token, so its individual letters are never directly visible to the model.
Do all models use the same tokenizer?
No. Each model family has its own vocabulary and method, so the same text gives different token counts on different models.
Conclusion
Tokenization is the first thing that happens to your prompt and the last thing that happens to the answer. It is a compromise between words and characters that lets a model handle any text with a manageable vocabulary. Knowing how it works explains token limits and bills, and why a system that writes fluent essays can still miscount the letters in a word.
Related articles
- How Large Language Models Predict the Next Word
- How Embeddings Turn Words Into Numbers
- Why Unicode Exists (and Why Emojis Break Things)
- Transformers Explained: The Architecture Behind Modern AI
