Concept icon
Concept

Mixture-of-Experts (MoE) for Efficient Inference

Mixture-of-Experts (MoE) models exemplify an efficient architecture applicable to LLM inference. In this approach, different 'expert' sub-networks are placed on separate devices, and only the experts relevant to a given input are activated for computation. This selective execution significantly boosts computational efficiency without sacrificing model quality.

0

1

Concept icon
Updated 2026-09-07

Contributors are:

Who are from:

Tags

Ch.5 Inference - Foundations of Large Language Models

Foundations of Large Language Models

Foundations of Large Language Models Course

Computing Sciences

Prep Sessions

Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.1 Edge Serving Bottlenecks and Dynamics - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Edge MoE Serving and Architectural Bottlenecks - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor