AI glossary
Context window
The maximum amount of text, measured in tokens, that a model can take into account at once: your messages, its replies, files and instructions.
Think of the context window as the model’s working memory for one conversation or request. Everything has to fit: the system prompt, the chat history, documents you attach, tool results and the answer being written. When it fills up, older content must be dropped or summarised.
Leading models in 2026 offer windows of around a million tokens — enough for several books. A bigger window does not guarantee the model uses every detail equally well, so put the most important information clearly and close to your question.
Example: Anthropic says that on its current models a million tokens is about 555,000 words: several books, or the code of a medium-sized project, fit in a single request.
In practice
- The window varies by model: Claude Haiku 4.5 has 200,000 tokens and Grok 4.7 has 500,000. They are all in our model comparison.
- In very long chats, starting a new conversation with a summary often gives better answers.
- You pay for every input token on each request, so a full window is expensive.
We explain it step by step in how LLMs work.


