Concept icon
Concept

Unified Slot Pool and Decode Seeding in Edge MoE Serving

In edge MoE serving, prefill double buffers and the decode expert cache are allocated from a single shared global slot pool rather than isolated memory partitions. Because both execution phases draw from the same residency pool, the serving engine eliminates the overhead of a dedicated prefill cache and avoids explicit phase-handoff reallocations. Moreover, expert weights already loaded in VRAM that persist across prefill automatically remain in the shared pool, effectively seeding the GPU expert cache for the subsequent latency-sensitive decode phase.

0

1

Concept icon
Updated 2026-09-07

Tags

Prep Sessions

Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.2 Pipelining and State Caching Mechanisms - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Pipelined Loading via Full-Layer Double Buffering - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor