Skip to content
estudIA

AI glossary

Pre-training

The first and most expensive stage of building a language model, in which it learns from a very large body of text by predicting the next token.

During pre-training the model reads a huge dataset — web pages, books, code and more — and learns to predict what comes next. This is where it acquires most of its general knowledge and language ability. It takes enormous computing resources.

Afterwards, post-training (instruction tuning, reinforcement learning from feedback and safety training) turns the raw model into a helpful assistant that follows instructions.

Example: Fresh from pretraining, a model can complete “The capital of France is…”, but if you ask it to “write me an apology email” it may just continue the text instead of doing it. Post-training teaches it to follow instructions.

In practice

  • What it learns in pretraining stops at the knowledge cutoff; for anything recent, give it sources.
  • Training data shapes the model: languages, topics and biases.
  • That is why models tend to do a little better in English than in languages with less text available.

Both phases, explained in how LLMs work.

Related terms

← Back to the glossary