Learn Before
How does semantic-aware LRU expert caching structure GPU VRAM allocation compared to traditional static expert placement?
0
1
Tags
Prep Sessions
Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Ch.3 Adaptive Runtime Policies and Device Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Bandwidth-Adaptive Decode and the q* Execution Policy - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
CUDA-Graph-Compatible Device-Side Cache Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Related
Which type of locality is demonstrated during MoE decoding when consecutive decoding steps within a layer repeatedly activate overlapping or recently utilized experts?
When a newly selected expert is not present in a full GPU VRAM cache during MoE decode, what two cache actions take place to manage expert residency?
How does semantic-aware LRU expert caching structure GPU VRAM allocation compared to traditional static expert placement?
During MoE decoding, what state change occurs within the caching system when an activated expert is already present in GPU VRAM (a cache hit)?
Bandwidth-Adaptive Miss Partitioning in MoE Decode