?
Problem

Sparse-Activation and Full-Expert-Pool Storage Mismatch in Edge MoE Serving

Mixture-of-Experts (MoE) models route each token through only a small subset of experts, with k≪Ek \ll E, reducing the active parameter footprint and per-token computation enough for frontier-scale execution to fit within consumer GPU memory. This sparsity does not proportionally reduce storage: the complete expert pool may still greatly exceed GPU capacity. Experts outside VRAM must therefore remain in host memory or persistent storage and be transferred to the GPU or executed on the CPU when routed. Edge MoE serving must reconcile a GPU-feasible active path with the cost of storing and accessing the full expert pool.

0

1

?
Updated 2026-09-12

Tags

Prep Sessions

Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.1 Edge Serving Bottlenecks and Dynamics - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Edge MoE Serving and Architectural Bottlenecks - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor