MoE vs Dense: how mixture-of-experts changes the memory math
A mixture-of-experts model needs the same memory to load as a dense model its size, but runs far faster. Here is the math, and where people get it wrong.
Servers Direct · May 20, 2026
A 235 billion parameter mixture-of-experts (MoE) model and a 235 billion parameter dense model use nearly the same memory to load. But the MoE model generates tokens many times faster. That speed difference is the entire reason mixture-of-experts models exist, and it is also where most people's hardware math goes wrong.
Typically people make the mistake that a mixture-of-experts (MoE) model is 'small' since only a couple billion of its total number of parameters are active at any given time. That is wrong. It is simply a large model that has an inexpensive runtime. If you understand this distinction you will know precisely what type of computer it will need.
What a mixture-of-experts model is

In a dense model, all of your parameters process every single input token. A dense model with 70 billion (70B) parameters performs 70 billion parameters worth of math for each token it reads or writes. Simple, predictable, and the reason large dense models are slow.
The Mixture of Experts (MoE) architecture alters a portion of this. Within each transformer block, the feed-forward layer has been decomposed into multiple, parallel versions known as experts. A small routing network examines each input token and selects the top few experts that will handle it. The remaining experts do not participate in processing that token. Attention layers typically remain dense and shared across all tokens, so only the feed-forward component is split into experts.
The gpt-oss-120b model demonstrates a clean way to achieve this. This model has 36 layers, 128 experts in each layer, and directs each token to the top four experts that will fire (OpenAI model card). Therefore, while there are a total of 116.8 billion possible parameters for all experts, approximately only 5.1 billion are active for any single input token. The routing process is performed by the router, not the user, and for each token and each layer, the router decides fresh which experts to activate.
Active parameters vs total parameters
The splitting of the MoE models provides a second count of parameters for each model. However, keeping track of these counts is the whole game.
- Total parameters are every expert plus the shared layers. This is what you have to store.
- Active parameters are the shared layers plus the handful of experts that fire for a given token. This is what actually computes.
On more recent models the name spells it out. Qwen3-235B-A22B means 235 billion total parameters, 22 billion active ('A' for active), and DeepSeek V3 is 671 billion total with 37 billion active per token. Model makers publish these splits directly, Mistral for example documented Mixtral's 46.7B total and 12.9B active at launch.
The memory math people get wrong

Here is the trap. Because gpt-oss-120b has only 5.1 billion parameters active (out of its 116.8 billion total), you might be tempted to estimate how much memory it needs based on what a 5-billion-parameter model would need. It does not work that way.
You must store each and every expert in your memory simultaneously, as the router can forward the subsequent token to any one of them. Routing occurs at run time for each token in each layer. As there is no mechanism to determine which experts can be excluded before runtime, you load each one. Like a dense model of that same number of parameters, the memory needed to load the model is defined by the total number of parameters.
gpt-oss-120b is post-trained in MXFP4, which stores a weight in about 4.25 bits. The MoE expert weights, which are over 90 percent of the parameter count, are quantized to MXFP4, while the rest (attention, embeddings) stay at higher precision. That brings the full 116.8-billion-parameter checkpoint to roughly 61 GB, which is what lets it fit on a single 80 GB GPU. The thing to notice is that the size tracks the total parameter count, not the active count: you store all 116.8 billion, not the 5.1 billion that fire per token. A dense model of the same total size at the same quantization would take about the same space to load.
Here is how the headline open models line up:
| Model | Total params | Active params | Experts / routing |
|---|---|---|---|
| Mixtral 8x7B | 46.7B | 12.9B | 8 experts, top-2 |
| gpt-oss-20b | 21B | 3.6B | 32 experts, top-4 |
| Mixtral 8x22B | 141B | 39B | 8 experts, top-2 |
| gpt-oss-120b | 116.8B | 5.1B | 128 experts, top-4 |
| Qwen3-235B-A22B | 235B | 22B | 128 experts, top-8 |
| DeepSeek V3 | 671B | 37B | 256 routed + 1 shared, top-8 on routed |
Notice that Mixtral 8x7B has an overall size of 46.7B total parameters, not 56B. The '8x7B' name suggests eight separate 7-billion-parameter models bolted together, but the attention layers are shared rather than duplicated, so the real total is lower. The active count, 12.9B, is also not 2x7B for the same reason. Names are marketing. The two numbers in the table are what matter.
DeepSeek V3 adds a wrinkle worth knowing. It keeps one expert always on (a shared expert) and routes each token to 8 of the other 256, rather than picking from a single flat pool. That always-on-plus-routed design is part of what separates newer DeepSeekMoE-style models from the original Mixtral approach, and you will see the "256 routed + 1 shared" phrasing in their docs.
Where the speedup actually comes from

What is the main advantage of the active count if memory capacity has to cover all parameters? Speed.
The computation per token grows with active parameters, not total. This matters most during decode, the token-by-token generation phase that is bound by memory bandwidth rather than raw compute (covered in DDR5 vs LPDDR5X). During decode the engine reads only the active experts' weights for each token. So a 235B-A22B model moves roughly 22 billion parameters worth of memory traffic per token, not 235 billion.
That is the mixture-of-experts bargain. You get the quality of a very large model (the knowledge of a 235B model) at the decode speed of a much smaller one (the speed of a 22B dense model), and you pay for it in memory capacity. It holds at that speed only as long as you have the memory to keep all 235 billion parameters resident.
Why this shapes the hardware you need
The memory-capacity-but-modest-bandwidth profile is exactly what unified-memory machines offer. A 128 GB Lenovo ThinkStation PGX can hold gpt-oss-120b at MXFP4 (about 62 GB) with room to spare, and it decodes at the speed of the 5.1 billion active parameters rather than the full 117 billion. A dense 117B model at the same quantization would also fit, but it would decode more than twenty times slower because every token would touch every parameter.
This is why mixture-of-experts models pair so well with capacity-rich local hardware. A 256 GB or 512 GB Mac Studio holds Qwen3-235B-A22B at 4-bit and runs it at 22B-active speed. The same logic extends up the line to multi-GPU GPU servers when total capacity exceeds what one box can hold, and down to discrete-GPU workstations for the smaller MoE models like Mixtral 8x7B that fit in 24 to 48 GB of VRAM.
The batch-size caveat
There is an honest limit to the low-cost-to-run story, and it shows up under load. At batch size 1, each token touches only its top-k experts, so memory traffic stays near the active-parameter count. But when you serve many users at once, different tokens in the batch route to different experts. Across a large enough batch you end up touching most of the experts anyway, and the effective active parameters per batch climb toward the total. The per-token efficiency that makes MoE attractive for a single local user erodes as concurrency rises, which is why production MoE serving leans on expert parallelism and careful routing. MoE is cheapest exactly where local inference lives: low concurrency, one or a few users.
Frequently Asked Questions

Does a mixture-of-experts model use less memory than a dense model?
No. The same amount of memory is needed to load a sparse model as a dense model with the same total parameter count, since every expert has to be resident in case the router selects it. What MoE reduces is compute and memory bandwidth per token, not the capacity required to hold the weights.
What is the difference between total and active parameters?
Total parameters are everything the model contains and what you must store in memory. Active parameters are the shared layers plus the few experts that fire for a given token, and they determine how much computation and memory traffic each token costs. gpt-oss-120b has 116.8B total and 5.1B active.
Why does Qwen3-235B-A22B run faster than a 70B dense model?
Because it only activates 22 billion parameters per token, while a 70B dense model activates all 70 billion. Decode speed tracks active parameters, so the 235B MoE decodes faster than the 70B dense model, even though it holds far more total knowledge. The catch is that it needs enough memory to hold all 235 billion parameters.
Can I run a big MoE model on a unified-memory machine like a PGX or Mac Studio?
Yes, as long as the total parameters fit in the unified memory at your chosen quantization. This is one of the best use cases for unified-memory hardware: plenty of room to store the entire model, and because the MoE design only activates a small fraction of the weights per token, decode runs at the speed of the small active count rather than the large total.
Do all the experts need to be in memory at once?
Yes. The router selects experts per token at runtime, so any expert can be needed at any moment. Streaming experts in and out of memory from disk or system RAM is possible but slow enough that it defeats the point. In practice you size the machine to hold the whole model.
Is a mixture-of-experts model always faster than a dense model?
At low concurrency, yes, because each token activates only a fraction of the weights. At high concurrency the advantage shrinks, because a large batch of tokens collectively routes to most of the experts, pushing effective compute back toward the total parameter count. MoE is at its best for single-user and small-team local inference.
The one sentence to remember: a mixture-of-experts model costs total-parameter memory to hold and active-parameter compute to run. Size the memory for the big number, expect the speed of the small one.
If you are trying to fit a specific MoE model on local hardware, we can do the capacity math with you. Talk to an engineer and we will match the model to a machine that can hold it.
Spec a system based on this guide.
Tell us what you took from this article — we’ll come back with a configuration.
Need rack-scale or integrated infrastructure?
For multi-node clusters, integrated racks, and larger enterprise deployments, our parent company Equus Compute Solutions handles the build. Same engineering team and supply chain, scaled up.

