AI glossary
Pre-training
The first and most expensive stage of building a language model, in which it learns from a very large body of text by predicting the next token.
During pre-training the model reads a huge dataset — web pages, books, code and more — and learns to predict what comes next. This is where it acquires most of its general knowledge and language ability. It takes enormous computing resources.
Afterwards, post-training (instruction tuning, reinforcement learning from feedback and safety training) turns the raw model into a helpful assistant that follows instructions.
Example: Fresh from pretraining, a model can complete “The capital of France is…”, but if you ask it to “write me an apology email” it may just continue the text instead of doing it. Post-training teaches it to follow instructions.
In practice
- What it learns in pretraining stops at the knowledge cutoff; for anything recent, give it sources.
- Training data shapes the model: languages, topics and biases.
- That is why models tend to do a little better in English than in languages with less text available.
Both phases, explained in how LLMs work.