?
Problem

Expert Transfer Bottleneck in Edge MoE Prefill

During prefill, a prompt supplies hundreds to thousands of tokens to each layer. Although each token routes to only kk experts, the union of selections across the prompt typically activates nearly the entire expert set. When the full model exceeds GPU memory, almost the complete expert pool must therefore be streamed over the CPU–GPU interconnect during each prefill pass. Transfer time consequently scales with the model's total parameter footprint rather than its sparse per-token active path, causing multi-second I/O latency and GPU idle time under on-demand expert loading.

0

1

?
Updated 2026-09-12

Tags

Prep Sessions

Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.1 Edge Serving Bottlenecks and Dynamics - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Prefill Transfer and Context Recomputation Challenges - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.2 Pipelining and State Caching Mechanisms - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Pipelined Loading via Full-Layer Double Buffering - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor