Explain why static expert placement leads to inefficiencies during the MoE decode phase, detailing where expert computations execute and how system hardware resources are impacted.
0
1
Tags
Prep Sessions
Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Ch.1 Edge Serving Bottlenecks and Dynamics - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Decode Cache Misses and Host Bandwidth Bottlenecks - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Related
Why do static expert placements capture only a small fraction of routed requests during the decode phase of MoE models?
In hybrid MoE serving systems, static expert allocation to GPU VRAM or host memory takes place dynamically during the decode phase.
Under static expert placement during MoE decode, what state is the PCIe interconnect left in when expert computations miss the GPU cache?
Explain why static expert placement leads to inefficiencies during the MoE decode phase, detailing where expert computations execute and how system hardware resources are impacted.
Host DRAM Bandwidth Bottleneck in CPU Expert Execution