Learn Before
Semantic-Aware LRU Expert Caching in MoE Decode
During the decoding phase of Mixture-of-Experts (MoE) serving, token routing demonstrates strong temporal locality: consecutive decoding steps within an MoE layer repeatedly activate overlapping or recently utilized experts. Rather than freezing expert placement with static assignments determined at startup, semantic-aware expert caching maintains a shared Least Recently Used (LRU) residency space in GPU VRAM across all layers. As generation proceeds, a cache hit refreshes the recency of an expert, a cache fill admits newly selected experts into VRAM, and eviction removes the least recently demanded expert. This dynamic caching mechanism allows scarce GPU memory to continuously adapt to the model's evolving working set during autoregressive generation.
0
1
Tags
Prep Sessions
Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Ch.3 Adaptive Runtime Policies and Device Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Bandwidth-Adaptive Decode and the q* Execution Policy - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
CUDA-Graph-Compatible Device-Side Cache Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Related
Semantic-Aware LRU Expert Caching in MoE Decode
Bandwidth-Adaptive Miss Partitioning in MoE Decode
Residual Host Bandwidth Formulation for CPU Expert Execution
Optimal Expert Miss Split Ratio (q* Policy)
Hardware-Specific Trade-Off in Serving MoE Decode Cache Misses
Device-Side Dynamic Cache Control via Data-Represented CUDA Graphs
Single-Pass Top-K LRU Victim Selection
Fused Multi-Bank Expert Transfer via Device-Resident Work Lists
Graph-Resident Heterogeneous CPU-GPU Execution Replay
Semantic-Aware LRU Expert Caching in MoE Decode
Bandwidth-Adaptive Miss Partitioning in MoE Decode
Learn After
Which type of locality is demonstrated during MoE decoding when consecutive decoding steps within a layer repeatedly activate overlapping or recently utilized experts?
When a newly selected expert is not present in a full GPU VRAM cache during MoE decode, what two cache actions take place to manage expert residency?
How does semantic-aware LRU expert caching structure GPU VRAM allocation compared to traditional static expert placement?
During MoE decoding, what state change occurs within the caching system when an activated expert is already present in GPU VRAM (a cache hit)?
Bandwidth-Adaptive Miss Partitioning in MoE Decode