Normalized Expert Bank Representation
To streamline heterogeneous execution across hardware components, model-specific Mixture-of-Experts checkpoint layouts are normalized into a unified set of expert banks. Each bank organizes parameters using a flattened layer–expert index as its leading dimension, where is the layer index, is the expert index within that layer, and is the total number of routed experts per layer. Rows across all banks sharing the same flattened index collectively constitute one full expert. This uniform organization guarantees that both device-side GPU kernels and host-side CPU worker threads reference a shared logical expert identity, irrespective of the underlying tensor representations.
0
1
Tags
Prep Sessions
Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Ch.4 Platform Adaptation and Performance Evaluation - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Expert Storage Formats and Platform Adaptation - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Learn After
In a normalized expert bank representation, which formula defines the leading dimension index for an expert with index in layer , given routed experts per layer?
According to the normalized expert bank structure, what collectively constitutes one full expert across the unified banks?
FreeToken Weight (FTW) Storage Format