Causation

Host DRAM Bandwidth Bottleneck in CPU Expert Execution

At small batch sizes typical of edge decoding, Mixture-of-Experts (MoE) expert computation is memory-bandwidth-bound because each token requires streaming the full weight tensors of its routed experts once. Consumer CPUs connected through dual-channel DRAM have peak memory bandwidths of approximately 50 GB/s50\text{ GB/s} with DDR4 and 80–90 GB/s80\text{--}90\text{ GB/s} with DDR5, whereas modern discrete GPUs achieve 1.0–1.8 TB/s1.0\text{--}1.8\text{ TB/s} from on-package VRAM. This host-memory bandwidth constraint limits the throughput of missed experts executed entirely on the CPU to a small fraction of the GPU's potential, regardless of the number of available CPU cores.

0

1

Updated 2026-09-12

Tags

Prep Sessions

Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.1 Edge Serving Bottlenecks and Dynamics - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Decode Cache Misses and Host Bandwidth Bottlenecks - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor