Concept icon
Concept

Host DRAM Bandwidth Bottleneck in CPU Expert Execution

At small batch sizes typical of edge decoding, MoE expert computation is memory-bandwidth bound because each token requires streaming the full weight tensors of its routed experts once. Consumer desktop and laptop platforms connect CPUs through dual-channel DRAM, offering peak memory bandwidths of approximately 50 GB/s50\text{ GB/s} for dual-channel DDR4 and 8090 GB/s80\text{--}90\text{ GB/s} for DDR5. In contrast, modern discrete GPUs achieve 1.01.8 TB/s1.0\text{--}1.8\text{ TB/s} from on-package VRAM. Due to this severe host memory bandwidth bottleneck, executing missed experts purely on the CPU caps decode throughput at a small fraction of the GPU's potential, regardless of the number of available CPU cores.

0

1

Concept icon
Updated 2026-09-07

Tags

Prep Sessions

Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.1 Edge Serving Bottlenecks and Dynamics - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Decode Cache Misses and Host Bandwidth Bottlenecks - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor