Concept icon
Concept

Hardware-Specific Trade-Off in Serving MoE Decode Cache Misses

When an active expert is not present in GPU VRAM during decoding, the system can either transfer its weights across PCIe to execute on the GPU, or compute it directly on the CPU where the weights reside in host RAM. Neither extreme is universally optimal: exclusively moving weights over PCIe leaves available host DRAM bandwidth and CPU cores idle if host memory can supply more data than the bus can move; conversely, computing entirely on the CPU leaves the PCIe interconnect idle and forfeits future cache hits that a VRAM cache fill provides. Because the balance between host memory bandwidth and PCIe transfer bandwidth varies widely across machines (such as LPDDR5 laptops versus high-speed PCIe desktops), the optimal division of miss handling cannot be statically predetermined and must be tuned to specific hardware resources.

0

1

Concept icon
Updated 2026-09-07

Tags

Prep Sessions

Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.1 Edge Serving Bottlenecks and Dynamics - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Decode Cache Misses and Host Bandwidth Bottlenecks - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.3 Adaptive Runtime Policies and Device Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Bandwidth-Adaptive Decode and the q* Execution Policy - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor