Concept icon
Concept

Dynamic VRAM Availability and Memory Split Shifts in Edge MoE Serving

Unlike datacenter environments with dedicated hardware, edge devices share the GPU among concurrent consumer applications such as the desktop compositor, web browsers, and games, which dynamically claim or release gigabytes of VRAM. Consequently, the total VRAM available to an MoE serving engine fluctuates across launches and during execution. Furthermore, the optimal internal allocation of GPU memory shifts over time: as multi-turn agentic workloads accumulate conversational context, Key-Value (KV) cache demand grows substantially while the MoE expert working set remains relatively stable. A static division chosen at session start becomes mismatched over time, requiring serving systems to elastically adjust both their total VRAM footprint and its internal split between KV cache pages and expert slots at runtime without restarting the engine.

0

1

Concept icon
Updated 2026-09-07

Tags

Prep Sessions

Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.1 Edge Serving Bottlenecks and Dynamics - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Non-Dedicated Edge Resource Dynamics - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.3 Adaptive Runtime Policies and Device Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Elastic Runtime Cache Reconfiguration and Fast Bootstrap - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor