AI glossary
Quantization
Storing a model’s weights with fewer bits so it uses less memory and runs faster, at a small cost in quality.
Weights are usually trained as 16- or 32-bit numbers. Quantization converts them to 8, 4 or even fewer bits. A model that needed a data-centre GPU can then fit on a laptop or phone.
It is especially common for running open-weights models locally. Aggressive quantization saves more memory but degrades answers more, so test the trade-off for your use.
Example: An 8-billion-parameter model takes about 16 GB at 16 bits per weight. Quantized to 4 bits it drops to about 4 to 5 GB and fits on a laptop with little memory.
In practice
- For local models, 4 or 8 bits is usually a good balance; below that, quality drops faster.
- Compare the quantized version with the original on your own tasks.
- If you use an API, the provider handles this: it only matters when you run models on your own machine.