MoE: Total vs. Active Parameters

黎 浩然/ 11 10 月, 2026/ 大语言模型/LARGELANGUAGEMODEL/LLM, 机器学习/MACHINELEARNING, 研究生/POSTGRADUATE, 计算机/COMPUTER/ 0 comments

When an MoE model lists “total” and “active” parameters, ask how they are counted. Total parameters describe stored model weights. Active parameters usually describe the subset involved in processing one token. A smaller active count does not mean the other experts can be deleted, or that speed matches a dense model of the same active size.

Contents
  1. Count an invented model first
  2. Sparse invocation does not erase experts
  3. One batch can overload one expert
  4. A parameter ratio is not a speed ratio
  5. Keep deployment comparisons controlled
  6. Sources

Count an invented model first

Suppose there are 200 million shared parameters and eight experts, each holding 100 million distinct parameters. The router selects two experts per token, all shared parameters participate in this path, and the router itself is included in the shared count:

Total = 200 M + 8 × 100 M = 1,000 M
Involved per token = 200 M + 2 × 100 M = 400 M
Stored parameters: 1,000 M

Used per token: 400 M

Original parameter-count comparison for an invented architecture: 200 M shared, eight 100 M experts, two selected per token. These bars measure parameter counts, not speed or measured VRAM.

This simplified example counts one expert set and omits subtleties of reused-parameter accounting. Real architectures may have several expert layers, shared experts or unequal modules. Sum the actual components and read the model’s counting convention instead of applying this formula universally.

Storing all one billion parameters tightly at 16 bits gives 2×10⁹ bytes, or 2 GB of raw weights. The 400 million per-token count does not reduce the complete weight budget to 0.8 GB. These figures exclude metadata, caches and runtime workspaces.

Sparse invocation does not erase experts

A sparse MoE layer routes a token representation to selected experts. Different tokens can choose different experts, and a whole batch can touch every expert. Weights unused by one token remain part of the model.

The Switch Transformer preprint appeared in 2021 and the peer-reviewed JMLR article in 2022. Its Switch routing selects one expert per token. The top-two example here is original counting arithmetic, not a reproduction of that paper’s configuration or experiments.

Deployment may distribute experts across devices or use other storage arrangements. Whether every weight is locally resident depends on the service. Moving weights elsewhere does not remove them; reading, transfer and scheduling costs still matter.

One batch can overload one expert

The following invented routes assign two experts to each of eight tokens. There are 8×2=16 assignments, averaging two per expert, yet the busiest expert receives seven. Every expert is touched. A small per-token active set and a balanced batch are different properties.

from collections import Counter

shared, per_expert, experts, top_k = 200, 100, 8, 2  # millions
assert 1 <= top_k <= experts
total = shared + experts * per_expert
active = shared + top_k * per_expert
assert (total, active) == (1000, 400)
print("total / active per token:", total, active, "million")

# Invented routing decisions: eight tokens, two experts per token.
routes = [(0,1), (0,2), (0,3), (0,4),
          (0,5), (0,6), (0,7), (1,2)]
load = Counter(e for route in routes for e in route)
counts = [load[e] for e in range(experts)]
assert sum(counts) == len(routes) * top_k == 16
assert counts == [7,2,2,1,1,1,1,1]
assert len(load) == experts
print("assignments per expert:", counts)
print("mean / peak:", sum(counts)/experts, max(counts))

The code was executed, checking total and per-token counts, assignment totals, individual loads and expert coverage. It learns no router, invokes no LLM and measures no time. One assignment also does not imply equal GPU work.

Another ideal batch could give each expert exactly two assignments. Both batches would have the same total work count but different bottlenecks. Grouping, batched matrix multiplication, communication and device speed make average load insufficient for predicting real latency.

A parameter ratio is not a speed ratio

The example’s per-token parameter count is 40% of the total, but this establishes no 40%-of-time inference claim. Shared computation, attention, routing, transfers and expert computation all contribute. Expert parallelism can also make one device wait for another.

Number What it describes What it does not establish alone
Total parameters Base scale of weight storage Peak inference memory
Active parameters Per-token weight subset under a stated convention End-to-end latency or quality
Expert assignments Count of routed tasks Actual communication and computation time

The Switch paper discusses communication, load and training stability. No reported speedup is reproduced here. Results under particular training and hardware conditions are not promises for arbitrary MoE inference services.

Keep deployment comparisons controlled

Record counting conventions, expert count, experts selected per token, weight precision and device placement. Test comparable input/output lengths and concurrency, measuring peak memory, completed throughput, time to first token and later generation intervals while retaining task-quality results.

Quantization bit width changes weight representation. GQA changes KV sharing. MoE concerns which experts a token uses. Separate these mechanisms rather than replacing a complete deployment budget with one smaller active-parameter number.

中文版

Sources

Switch Transformers, JMLR 2022; Peer-reviewed paper; arXiv submission history.

Support

If this article helped you, you can support this site.

WeChat support QR code; click to enlarge
WeChat
Alipay support QR code; click to enlarge
Alipay
Buy Me a Coffee; support this site
Buy Me a Coffee

Click a QR code to enlarge. More options: support page。

Share this Post

Leave a Comment

您的邮箱地址不会被公开。 必填项已用 * 标注

*
*