Sparse-Activation and Full-Expert-Pool Storage Mismatch in Edge MoE Serving
Mixture-of-Experts (MoE) models route each token through only a small subset of experts, with , reducing the active parameter footprint and per-token computation enough for frontier-scale execution to fit within consumer GPU memory. This sparsity does not proportionally reduce storage: the complete expert pool may still greatly exceed GPU capacity. Experts outside VRAM must therefore remain in host memory or persistent storage and be transferred to the GPU or executed on the CPU when routed. Edge MoE serving must reconcile a GPU-feasible active path with the cost of storing and accessing the full expert pool.
0
1
Tags
Prep Sessions
Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Ch.1 Edge Serving Bottlenecks and Dynamics - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Edge MoE Serving and Architectural Bottlenecks - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Related
Mixture-of-Experts (MoE) for Efficient Inference
Sparse-Activation and Full-Expert-Pool Storage Mismatch in Edge MoE Serving
Experts as Modular FFNs in LLM MoE Models
A large language model is deployed for inference across 8 powerful processing units. In one configuration, the entire model's computational graph is activated across all 8 units for every input. In a second configuration, the model is structured with 8 distinct 'expert' sub-networks, one on each unit. For a given input, a routing mechanism selects only the 2 most relevant expert sub-networks to perform computations. What is the primary efficiency benefit of the second configuration for processin
In a Mixture-of-Experts (MoE) architecture, all expert sub-networks must be hosted on a single hardware device.
What effect does selective execution have on computational efficiency and model quality in MoE inference?
Based on MoE operational principles, which expert sub-networks are activated for computation on this request?
Sparse-Activation and Full-Expert-Pool Storage Mismatch in Edge MoE Serving
Learn After
How does a Mixture-of-Experts (MoE) architecture allow the computation of a frontier-scale model to fit within consumer GPU memory on an edge device?
What constitutes the central systems challenge of serving Mixture-of-Experts (MoE) models on edge devices?
What is the impact of architectural sparsity on the memory required to store an MoE model's full expert pool?
Where do inactive experts reside when the complete parameter footprint of an MoE model exceeds the memory of an edge device's GPU?
Match each MoE serving parameter or edge concept to its corresponding description.
Identify the primary systems bottleneck causing latency degradation in this scenario, and explain why expanding the total expert pool () while keeping the number of active experts () constant worsens this bottleneck.
In an edge MoE deployment where the full model footprint exceeds GPU capacity, by what mechanism do inactive experts participate in token processing?
In edge Mixture-of-Experts (MoE) serving, consumer GPU memory only needs to accommodate the active compute footprint for a given token rather than the entire parameter footprint simultaneously during execution.
Which two operational metrics are drastically reduced when an MoE architecture routes each token to a sparse subset of experts ()?