Learn Before
Relation

Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Learners will gain a comprehensive understanding of the systems architecture and runtime techniques required to serve frontier-scale Mixture-of-Experts (MoE) models on consumer and workstation hardware. The materials explore prefill pipelining, semantic-aware recurrent state reuse, bandwidth-adaptive hybrid CPU-GPU execution, and CUDA graph integration. Learners will develop the ability to analyze and optimize memory, interconnect, and compute trade-offs across dynamic, resource-constrained edge environments.

0

1

Updated 2026-09-07

Tags

Prep Sessions