Skip to content

What is Mixture of Experts (MoE)?

Generative AI & LLMs, explained by the engineers who build it. Definition, how it works, use cases and common questions.

MoE definition

Mixture of experts (MoE) is a neural network architecture in which a model contains many specialized subnetworks, called experts, and a router activates only a few of them for each input token. This lets a model have a very large total number of parameters while using only a fraction for each prediction, improving quality without a proportional increase in compute.

How mixture of experts works

In a standard dense transformer, every parameter takes part in processing every token. In an MoE transformer, some layers, usually the feed-forward layers, are replaced by a set of experts: for example 8, 64 or more parallel feed-forward networks. A small gating network, the router, looks at each token and selects the top one or two experts, or a few more in fine-grained designs, to process it, then combines their outputs weighted by its scores.

Because only the selected experts run, compute per token depends on the active parameters rather than the total. A model can hold the capacity of a very large network while running roughly as fast as a much smaller one. The idea dates back to research in the early 1990s and was scaled up for language models by Google researchers from the late 2010s onward.

Examples of MoE models

Mistral AI's Mixtral 8x7B, released in 2023, brought MoE to widely used open-weight models, with eight experts per layer and two active per token. DeepSeek's V3 and R1 models then used fine-grained experts plus shared experts to reach strong performance at lower training cost. MoE has since become the standard design for flagship open-weight models: recent DeepSeek, Qwen, Kimi, GLM, gpt-oss and Llama releases all include MoE models, often alongside smaller dense variants.

Several leading proprietary models are reported or confirmed to use MoE designs as well, although providers often do not publish architecture details. For buyers, architecture matters less than measured quality, cost and latency on their own tasks, and those are what should drive model selection.

Benefits and trade-offs

MoE changes the economics of large models, but it moves complexity elsewhere rather than removing it entirely, mostly into memory and serving infrastructure. The main benefits and costs for teams training or serving these models in production are:

  • Benefit: more model capacity for the same compute per token, often improving quality per unit of training and inference cost
  • Benefit: faster inference than a dense model of the same total size
  • Cost: all experts must sit in memory, so GPU memory needs follow total parameters, not active ones
  • Cost: routing adds complexity, including load balancing so some experts are not overused while others sit idle
  • Cost: distributed serving across GPUs needs fast interconnects and specialized inference engines
  • Cost: fine-tuning can be less stable than with dense models

What MoE means for businesses using AI

Most teams meet MoE indirectly, through APIs whose price and speed reflect these efficiencies. It becomes a practical question when self-hosting open models: an MoE model with a modest active parameter count may still need several GPUs because of its total size, while a dense model of similar quality might fit on one. Quantization reduces memory but does not change the basic trade-off.

When choosing a model, compare quality on your own tasks, memory footprint, throughput and cost per request rather than parameter counts. Nexzem benchmarks dense and MoE open models alongside hosted APIs during LLM development, and plans inference infrastructure around the results.

MoE: common questions

Something else on your mind? Ask a consultant and get a reply within one business day.

Is a mixture of experts model smarter than a dense model?

Not inherently. MoE is an efficiency technique: for a given compute budget, it lets a model have more parameters, which often improves quality. A dense model with the same total parameters may perform similarly or better but costs far more to run. Compare models on benchmarks and your own tasks, not architecture labels.

What does 8x7B mean in a model name?

It indicates eight experts built on a 7-billion-parameter base design. Because only the expert layers are duplicated while attention and other layers are shared, the total is well below 56 billion parameters, and only two experts are active per token, so compute per token is much lower than the total suggests.

Do experts specialize in topics like math or law?

Not in the way the name suggests. Research on trained MoE models finds that experts often specialize in token-level patterns, such as syntax, punctuation or particular kinds of words, rather than clean human topics. The router learns whatever split reduces training loss, which rarely maps neatly onto subject areas.

Keep exploring the generative ai & llms glossary

Need MoE in your product?

A solutions consultant replies within one business day with next steps, a rough estimate and a suggested team.