Activity (Process)

Pipelined Full-Layer Double Buffering in MoE Prefill

To hide massive expert weight movement behind computation during the prefill phase, serving systems can employ full-layer double buffering. Because prompt processing activates nearly the entire expert set across each layer, the system allocates two full-layer buffers in GPU memory from the global slot pool rather than fetching individual experts on demand. While the GPU computes the routed experts for layer ll using the active buffer, a dedicated transfer stream concurrently pre-fetches the entire expert set of layer l+1l + 1 over PCIe into the alternate buffer. Loading the complete layer allows transfer to begin immediately in the background before the router determines token routing for that layer. Once layer computation completes, the buffers swap roles. If GPU memory is insufficient to allocate two full layers, the engine falls back to on-demand expert loading to prevent memory oversubscription.

0

1

Updated 2026-09-07

Tags

Prep Sessions

Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.2 Pipelining and State Caching Mechanisms - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Pipelined Loading via Full-Layer Double Buffering - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.4 Platform Adaptation and Performance Evaluation - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Agentic Workload Serving and Cross-Hardware Performance - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Related