Concept icon
Concept

Expert Transfer Bottleneck in Edge MoE Prefill

During the prefill phase of inference, processing a prompt involves hundreds to thousands of tokens per layer. Although each individual token routes to only a sparse subset of kk experts, the union of routed tokens across the entire prompt typically activates nearly the entire expert set in every layer. In an edge deployment where the full model exceeds available GPU memory (VRAM), almost the complete expert pool must be streamed over the CPU–GPU interconnect (e.g., PCIe) during each prefill pass. Consequently, prefill transfer time scales with the total parameter footprint of the model rather than the sparse active path, introducing significant multi-second I/O latency and leaving the GPU idle while weights are fetched on demand.

0

1

Concept icon
Updated 2026-09-07

Tags

Prep Sessions

Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.1 Edge Serving Bottlenecks and Dynamics - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Prefill Transfer and Context Recomputation Challenges - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.2 Pipelining and State Caching Mechanisms - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Pipelined Loading via Full-Layer Double Buffering - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor