AI glossary
Transformer
The neural network architecture behind modern language models. Its “attention” mechanism lets every token look at every other token in the input.
Introduced in the 2017 paper Attention Is All You Need, the transformer replaced older designs that read text one word at a time. With attention, the model weighs how relevant each part of the input is to every other part, all in parallel. That makes it good at long-range connections — like linking a pronoun to the noun it refers to three paragraphs earlier — and efficient to train on GPUs.
The “T” in GPT stands for transformer.
Example: In the sentence “The dog did not cross the road because it was too tired”, attention helps the model see that “it” refers to the dog, not the road.
In practice
- Almost all current language models, and many image and audio models, are built on it.
- The cost of attention grows with the length of the text: one reason long contexts are expensive.
- To understand it in depth, our resources list free courses.
How it fits into the bigger picture: how LLMs work.