Activity (Process)

Runtime Cache Reconfiguration at Scheduler Safe Points

To adapt to shifting VRAM availability and growing Key-Value (KV) cache demands during execution, the serving engine dynamically reconfigures its GPU memory allocation without process restarts. After dedicating VRAM to non-expert weights and fixed runtime structures, the system partitions the remaining memory between KV cache pages and complete-expert cache slots. At any scheduler safe point, the runtime can rebuild the GPU expert cache to match a revised VRAM budget and re-establish its static execution graph. Because the CPU-resident expert pool acts as the permanent source of truth for all expert weights, altering or flushing the GPU expert cache affects only inference performance rather than computation correctness.

0

1

Updated 2026-09-07

Tags

Prep Sessions

Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.3 Adaptive Runtime Policies and Device Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Elastic Runtime Cache Reconfiguration and Fast Bootstrap - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor