MoE: Total vs. Active Parameters
When an MoE model lists “total” and “active” parameters, ask how they are counted. Total parameters describe stored model weights. Active parameters usually describe the subset involved in processing one token. A smaller active count does not mean the other experts can be deleted, or that speed matches a dense model of the same active size.
Contents
Count an invented model first
Suppose there are 200 million shared parameters and eight experts, each holding 100 million distinct parameters. The router selects two experts per token, all shared parameters participate in this path, and the router itself is included in the shared count:
Total = 200 M + 8 × 100 M = 1,000 M
Involved per token = 200 M + 2 × 100 M = 400 M
This simplified example counts one expert set and omits subtleties of reused-parameter accounting. Real architectures may have several expert layers, shared experts or unequal modules. Sum the actual components and read the model’s counting convention instead of applying this formula universally.
Storing all one billion parameters tightly at 16 bits gives 2×10⁹ bytes, or 2 GB of raw weights. The 400 million per-token count does not reduce the complete weight budget to 0.8 GB. These figures exclude metadata, caches and runtime workspaces.
Sparse invocation does not erase experts
A sparse MoE layer routes a token representation to selected experts. Different tokens can choose different experts, and a whole batch can touch every expert. Weights unused by one token remain part of the model.
The Switch Transformer preprint appeared in 2021 and the peer-reviewed JMLR article in 2022. Its Switch routing selects one expert per token. The top-two example here is original counting arithmetic, not a reproduction of that paper’s configuration or experiments.
Deployment may distribute experts across devices or use other storage arrangements. Whether every weight is locally resident depends on the service. Moving weights elsewhere does not remove them; reading, transfer and scheduling costs still matter.
One batch can overload one expert
The following invented routes assign two experts to each of eight tokens. There are 8×2=16 assignments, averaging two per expert, yet the busiest expert receives seven. Every expert is touched. A small per-token active set and a balanced batch are different properties.
from collections import Counter
shared, per_expert, experts, top_k = 200, 100, 8, 2 # millions
assert 1 <= top_k <= experts
total = shared + experts * per_expert
active = shared + top_k * per_expert
assert (total, active) == (1000, 400)
print("total / active per token:", total, active, "million")
# Invented routing decisions: eight tokens, two experts per token.
routes = [(0,1), (0,2), (0,3), (0,4),
(0,5), (0,6), (0,7), (1,2)]
load = Counter(e for route in routes for e in route)
counts = [load[e] for e in range(experts)]
assert sum(counts) == len(routes) * top_k == 16
assert counts == [7,2,2,1,1,1,1,1]
assert len(load) == experts
print("assignments per expert:", counts)
print("mean / peak:", sum(counts)/experts, max(counts))
The code was executed, checking total and per-token counts, assignment totals, individual loads and expert coverage. It learns no router, invokes no LLM and measures no time. One assignment also does not imply equal GPU work.
Another ideal batch could give each expert exactly two assignments. Both batches would have the same total work count but different bottlenecks. Grouping, batched matrix multiplication, communication and device speed make average load insufficient for predicting real latency.
A parameter ratio is not a speed ratio
The example’s per-token parameter count is 40% of the total, but this establishes no 40%-of-time inference claim. Shared computation, attention, routing, transfers and expert computation all contribute. Expert parallelism can also make one device wait for another.
| Number | What it describes | What it does not establish alone |
|---|---|---|
| Total parameters | Base scale of weight storage | Peak inference memory |
| Active parameters | Per-token weight subset under a stated convention | End-to-end latency or quality |
| Expert assignments | Count of routed tasks | Actual communication and computation time |
The Switch paper discusses communication, load and training stability. No reported speedup is reproduced here. Results under particular training and hardware conditions are not promises for arbitrary MoE inference services.
Keep deployment comparisons controlled
Record counting conventions, expert count, experts selected per token, weight precision and device placement. Test comparable input/output lengths and concurrency, measuring peak memory, completed throughput, time to first token and later generation intervals while retaining task-quality results.
Quantization bit width changes weight representation. GQA changes KV sharing. MoE concerns which experts a token uses. Separate these mechanisms rather than replacing a complete deployment budget with one smaller active-parameter number.
Sources
Switch Transformers, JMLR 2022; Peer-reviewed paper; arXiv submission history.
Support
If this article helped you, you can support this site.
Click a QR code to enlarge. More options: support page。


