Concept icon
Concept

Device-Side Dynamic Cache Control via Data-Represented CUDA Graphs

In Mixture-of-Experts serving, expert caching is inherently dynamic because missing expert identities, transfer batch counts, and eviction slots vary at every generation step. Managing these decisions on the host CPU would force synchronous CPU–GPU device synchronizations at each MoE layer, destroying execution pipelining. To eliminate host intervention while retaining compatibility with static CUDA Graph capture, all routing-dependent decisions are executed entirely on the GPU and encoded as device data within fixed-shape buffers and device-resident valid counts. A dedicated device kernel deduplicates routed expert IDs, checks residency status, computes the bandwidth-proportional fetch quota qq, marks eviction victims, and maps logical routed identifiers into physical VRAM slots or CPU-execution flags without synchronizing back to the host.

0

1

Concept icon
Updated 2026-09-07

Tags

Prep Sessions

Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.3 Adaptive Runtime Policies and Device Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

CUDA-Graph-Compatible Device-Side Cache Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor