01.
The Active vs Total Parameter Paradox
MoE sparse routing only activates a subset of weights (e.g. 23B out of 501B) during the forward pass. However, all experts must remain resident in VRAM to avoid catastrophic PCIe swapping latency (100x slowdown). Active params determine token generation speed, but TOTAL params determine your physical GPU count.
02.
1-Million Context KV Cache Explosion
At 1,000,000 tokens, traditional Multi-Head Attention (MHA) KV caches require over 120 GB of memory PER concurrent user in FP16. Always enable FP8 KV caching (--kv-cache-dtype fp8) or verify that the model implements Multi-Head Latent Attention (MLA) before exposing 1M endpoints.
03.
PCIe vs NVLink All-Reduce Wall
Tensor Parallelism (TP) across consumer GPUs without NVLink (e.g. 4x RTX 4090 connected via PCIe Gen4 x16) suffers from severe All-Reduce communication bottlenecks. For MoE models above 100B, SXM5 NVLink interconnect (900 GB/s) is mandatory to prevent throughput collapse.