Optimal MoE Decode Cache-Fill Count ( Policy)
For missing experts of size , let experts use the concurrent PCIe cache-fill branch and execute on the CPU. When , their approximate execution times are and . Balancing the branches gives and therefore . The runtime rounds to an integer and retains at least one cache fill so that the cache continues warming. As approaches from above, approaches , reducing the policy to pure on-demand cache fill.
0
1
Tags
Prep Sessions
Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Ch.3 Adaptive Runtime Policies and Device Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Bandwidth-Adaptive Decode and the q* Execution Policy - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Related
Order the steps to determine and utilize residual host bandwidth for CPU expert execution under a saturated PCIe bus.
Calculate the residual host bandwidth (B_R) available for concurrent in-place CPU computation, and explain whether the CPU can execute missed experts concurrently without contending with or throttling the PCIe bus.
Optimal MoE Decode Cache-Fill Count ( Policy)
Learn After
A system encounters missing experts during decode, with PCIe bandwidth and host bandwidth . Based on the optimal execution policy, how many missing experts should be routed to the concurrent PCIe cache-fill branch?
In the runtime implementation of the execution policy, a minimum of at least one cache fill is retained to ensure continuous cache warming.
Under the execution policy, what behavior does the system dynamically degenerate to as the host bandwidth approaches PCIe bandwidth ()?
Explain the mathematical derivation used to determine the optimal expert cache-fill count for serving missing experts of size . Include the expressions for both concurrent branches, how they are equated, and the resulting closed-form policy.