Learn Before
Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Learners will gain a comprehensive understanding of the systems architecture and runtime techniques required to serve frontier-scale Mixture-of-Experts (MoE) models on consumer and workstation hardware. The materials explore prefill pipelining, semantic-aware recurrent state reuse, bandwidth-adaptive hybrid CPU-GPU execution, and CUDA graph integration. Learners will develop the ability to analyze and optimize memory, interconnect, and compute trade-offs across dynamic, resource-constrained edge environments.
0
1
Tags
Prep Sessions
Related
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Learn After
Ch.1 Edge Serving Bottlenecks and Dynamics - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
2608.16157v1.pdf
Ch.2 Pipelining and State Caching Mechanisms - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Ch.3 Adaptive Runtime Policies and Device Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Ch.4 Platform Adaptation and Performance Evaluation - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor