Skip to content
estudIA

AI glossary

Quantization

Storing a model’s weights with fewer bits so it uses less memory and runs faster, at a small cost in quality.

Weights are usually trained as 16- or 32-bit numbers. Quantization converts them to 8, 4 or even fewer bits. A model that needed a data-centre GPU can then fit on a laptop or phone.

It is especially common for running open-weights models locally. Aggressive quantization saves more memory but degrades answers more, so test the trade-off for your use.

Example: An 8-billion-parameter model takes about 16 GB at 16 bits per weight. Quantized to 4 bits it drops to about 4 to 5 GB and fits on a laptop with little memory.

In practice

  • For local models, 4 or 8 bits is usually a good balance; below that, quality drops faster.
  • Compare the quantized version with the original on your own tasks.
  • If you use an API, the provider handles this: it only matters when you run models on your own machine.

Related terms

← Back to the glossary