Concept icon
Concept

Semantic-Aware LRU Expert Caching in MoE Decode

During the decoding phase of Mixture-of-Experts (MoE) serving, token routing demonstrates strong temporal locality: consecutive decoding steps within an MoE layer repeatedly activate overlapping or recently utilized experts. Rather than freezing expert placement with static assignments determined at startup, semantic-aware expert caching maintains a shared Least Recently Used (LRU) residency space in GPU VRAM across all layers. As generation proceeds, a cache hit refreshes the recency of an expert, a cache fill admits newly selected experts into VRAM, and eviction removes the least recently demanded expert. This dynamic caching mechanism allows scarce GPU memory to continuously adapt to the model's evolving working set during autoregressive generation.

0

1

Concept icon
Updated 2026-09-07

Tags

Prep Sessions

Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.3 Adaptive Runtime Policies and Device Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Bandwidth-Adaptive Decode and the q* Execution Policy - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

CUDA-Graph-Compatible Device-Side Cache Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor