Concept icon
Concept

Bandwidth-Adaptive Miss Partitioning in MoE Decode

When decode-time cache misses occur, bandwidth-adaptive execution splits the set of mm missing experts MM into two disjoint subsets: a cache-fill set FF containing q=Fq = |F| experts, and a CPU-execution set CC containing mqm - q experts, such that M=FCM = F \cup C. Experts assigned to FF are transferred over the PCIe bus into GPU cache slots, evaluated on the GPU, and retained in VRAM for future reuse. Concurrently, experts assigned to CC execute in place directly from host DRAM by the CPU without modifying GPU cache residency. This concurrent partitioning utilizes the PCIe bus at full bandwidth while simultaneously using residual host memory bandwidth for CPU computation, avoiding GPU idle time and maintaining continuous cache updates.

0

1

Concept icon
Updated 2026-09-07

Tags

Prep Sessions

Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.3 Adaptive Runtime Policies and Device Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Bandwidth-Adaptive Decode and the q* Execution Policy - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

CUDA-Graph-Compatible Device-Side Cache Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor