Concept icon
Concept

Normalized Expert Bank Representation

To streamline heterogeneous execution across hardware components, model-specific Mixture-of-Experts checkpoint layouts are normalized into a unified set of expert banks. Each bank organizes parameters using a flattened layer–expert index lE+elE + e as its leading dimension, where ll is the layer index, ee is the expert index within that layer, and EE is the total number of routed experts per layer. Rows across all banks sharing the same flattened index collectively constitute one full expert. This uniform organization guarantees that both device-side GPU kernels and host-side CPU worker threads reference a shared logical expert identity, irrespective of the underlying tensor representations.

0

1

Concept icon
Updated 2026-09-07

Tags

Prep Sessions

Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.4 Platform Adaptation and Performance Evaluation - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Expert Storage Formats and Platform Adaptation - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor