Why is managing expert caching decisions on the host CPU disadvantageous during Mixture-of-Experts (MoE) serving?
0
1
Tags
Prep Sessions
Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Ch.3 Adaptive Runtime Policies and Device Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
CUDA-Graph-Compatible Device-Side Cache Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Related
Match each device kernel task to its corresponding role in device-side dynamic cache control.
To avoid host intervention while supporting static CUDA Graphs, variable expert quantities are tracked on the GPU using device-resident ___ counts within fixed-shape buffers.
Why is managing expert caching decisions on the host CPU disadvantageous during Mixture-of-Experts (MoE) serving?
According to the device-side cache control mechanism, what two destinations or indicators can the dedicated device kernel map logical routed identifiers into?
Single-Pass Top-K LRU Victim Selection
Fused Multi-Bank Expert Transfer via Device-Resident Work Lists
Graph-Resident Heterogeneous CPU-GPU Execution Replay