Concept icon
Concept

Systems Challenge of Edge MoE Serving

Mixture-of-Experts (MoE) architectures make frontier-scale model execution feasible on edge devices because routing each token to only a sparse subset of experts (kEk \ll E) drastically reduces the active parameter footprint and per-token computation. For example, a model may activate only a small fraction of its total parameter count for any single token, allowing the active compute footprint to fit within consumer GPU memory. However, architectural sparsity reduces computation without proportionally shrinking the memory required to store the full expert pool. When the complete model parameter footprint greatly exceeds GPU memory, inactive experts must reside in CPU host memory or secondary storage and enter the execution path on demand. This discrepancy creates the central systems challenge of edge MoE serving: sparse activation makes on-device computation feasible, but storing and streaming the complete expert pool makes efficient serving difficult.

0

1

Concept icon
Updated 2026-09-07

Tags

Prep Sessions

Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.1 Edge Serving Bottlenecks and Dynamics - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Edge MoE Serving and Architectural Bottlenecks - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Prefill Transfer and Context Recomputation Challenges - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Decode Cache Misses and Host Bandwidth Bottlenecks - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Non-Dedicated Edge Resource Dynamics - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Related