Learn Before
How does the GPU expert cache transition to a warm state in a cold-cache serving runtime?
0
1
Tags
Prep Sessions
Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Ch.3 Adaptive Runtime Policies and Device Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Elastic Runtime Cache Reconfiguration and Fast Bootstrap - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Related
How does the GPU expert cache transition to a warm state in a cold-cache serving runtime?
In traditional serving runtimes, a dedicated warmup phase is used prior to processing user requests in order to prime GPU execution caches and allocators.
What major performance penalty do traditional serving runtimes impose on consumer edge devices by requiring a dedicated GPU warmup phase?
Describe how a cold-cache serving engine handles initial expert requests when starting with an unpopulated GPU cache, detailing the specific execution and transfer mechanisms used.