AI glossary
Token
The unit of text a language model reads and writes: a whole word, part of a word, a number or a punctuation mark.
Models do not see letters or words directly; a tokenizer splits text into tokens first. How much text fits in a token depends on the tokenizer: Anthropic says a million tokens was about 750,000 words on its earlier models and is about 555,000 on its current ones. Other languages, code and unusual words often need more tokens than English.
Tokens matter in practice because API prices, speed and the context window are all measured in them. Each model family has its own tokenizer, so the same text can be a different number of tokens on different models.
Example: “Unbelievable!” might be split into tokens like “Un”, “believ”, “able” and “!”.
In practice
- Spanish and other languages usually take a few more tokens than English: allow for it when estimating costs.
- APIs provide tools to count a request’s tokens before you send it.
- A price “per million tokens” refers to input and output separately.
Prices per million tokens for each model: model comparison.

