Start here4 min read
How AI Language Models Work: LLMs, Tokens and Context
A plain-language explanation of what large language models are, how they generate text, what tokens and context windows are, and why it matters to you.
When you type a question into ChatGPT, Claude or Gemini, a large language model (LLM) writes the answer. You do not need to know the maths to use these tools well, but a few core ideas will help you understand why they are so useful, why they sometimes fail, and how to get better results.
“AI” today usually means generative AI
Artificial intelligence is an old and broad field. A spam filter, a chess program and a photo app that recognises faces are all AI. What changed in recent years is generative AI: models that produce new text, images, audio or code instead of only sorting or predicting.
LLMs are the generative models specialised in language. They are behind chat assistants, writing tools and coding agents.
An LLM predicts the next piece of text
At its core, a language model does one thing: given some text, it predicts what is likely to come next. It produces one small piece, adds it to the text, and predicts again. Repeated thousands of times, that simple step produces whole paragraphs, summaries, translations or programs.
That sounds too simple to be useful, but to predict text well across billions of examples, a model has to absorb a lot: grammar, facts, styles, how arguments are built, how code is structured. The modern design that made this work at scale is the transformer, introduced in 2017, whose “attention” mechanism lets each word take every other word in the input into account.
How a model is trained
Building an LLM happens in two broad stages:
- Pre-training. The model reads an enormous amount of text — web pages, books, code — and learns to predict the next piece. This is where it gets its general knowledge. It is extremely expensive and happens once per model.
- Post-training. The raw model is then taught to follow instructions, be helpful and refuse harmful requests, using curated examples and feedback from people and other models.
Training stops at some point, so every model has a knowledge cutoff. It knows nothing that happened afterwards unless it can search the web or you give it the information.
Tokens: how models read text
Models do not read letters or whole words. A tokenizer splits text into tokens: whole words, parts of words, numbers or punctuation. How much text fits in a token depends on the tokenizer: Anthropic says a million tokens was about 750,000 words on its earlier models and is about 555,000 on its current ones. Spanish and other languages often need a few more tokens for the same idea, and so does code.
Tokens matter in practice because:
- API prices are charged per million tokens, separately for what you send (input) and what the model writes (output).
- Speed is measured in tokens per second.
- Limits, like the maximum length of an answer, are counted in tokens.
The context window: the model’s working memory
The context window is the maximum number of tokens a model can consider at once. Everything in a conversation has to fit inside it: hidden instructions from the app, your messages, the model’s replies, attached files and the results of any tools it uses.
Leading models in 2026 offer context windows of around a million tokens, enough for several long books. That is very useful for analysing large documents or codebases, but two cautions apply:
- A bigger window does not mean the model weighs every detail equally. Put the important information clearly, and ask your question after the material, not buried in the middle.
- Long contexts cost more and are slower, because you pay for every input token on every request.
Outside of the context window, the model has no memory of you. If an app “remembers” earlier chats, it is because the app saves notes and adds them to the context.
Why you get different answers to the same question
When choosing the next token, the model does not always pick the single most likely one. A bit of controlled randomness makes writing more natural and varied. That is why asking the same question twice can give two different phrasings — or occasionally two different conclusions. For tasks where consistency matters, give precise instructions and examples, and check the important results.
What LLMs are good and bad at
Good at: drafting and editing text, summarising, translating, explaining concepts at different levels, brainstorming, extracting information into a structured format, and writing and reviewing code.
Weak at, unless given help: recent events (knowledge cutoff), exact facts they were never shown, precise arithmetic on large numbers without a calculator tool, and knowing when they are wrong. Because they generate what is plausible, they can state false things with confidence. This is called a hallucination, and it is the subject of our next guide.
Key takeaways
- An LLM predicts text one token at a time, using patterns learned from huge amounts of writing.
- Tokens are the unit of cost, speed and limits.
- The context window is all the model can “see” at once; nothing outside it exists for the model.
- Treat answers as a strong first draft: verify facts, numbers and sources that matter.
Next, read about hallucinations and privacy — the two things every new user should understand before relying on AI.
Frequently asked questions
Does ChatGPT or Claude search the internet to answer?
Only if web search is turned on. Otherwise the model answers from what it learned during training, which stops at its knowledge cutoff date.
Does the model remember my previous chats?
Within one conversation it sees the whole history, up to the size of its context window. Between conversations it only remembers things if the app has a memory feature and you have it turned on.
Is an LLM the same as AI?
No. An LLM is one kind of AI system, specialised in language. AI also includes image generators, recommendation systems, speech recognition and much more.


