Skip to content
estudIA

AI glossary

Mixture of experts (MoE)

A model design that splits the network into many “experts” and activates only a few of them for each token, so a huge model runs at the cost of a smaller one.

In a mixture-of-experts model, a small router decides which experts should handle each token. Only those experts do the work, so most of the model’s parameters sit idle at any moment. That makes very large models cheaper and faster to run.

The trade-off is complexity: MoE models need more memory to hold all the experts, and training them so every expert learns something useful is tricky. Several open-weight models, including DeepSeek’s, use this design.

Example: An MoE model can have hundreds of billions of parameters in total yet activate only some of them for each token. That is why it answers faster and more cheaply than a “dense” model of the same total size.

In practice

  • Model pages show two figures: total parameters and active parameters per token.
  • To run it locally, memory depends on the total, even though only the active ones do the work.
  • If you use it through an API you do not have to do anything different: it is an internal detail.

Related terms

Learn more

← Back to the glossary