Formula

Optimal MoE Decode Cache-Fill Count (q∗q^* Policy)

For mm missing experts of size SS, let qq experts use the concurrent PCIe cache-fill branch and m−qm-q execute on the CPU. When BH>BPB_H>B_P, their approximate execution times are Tfill(q)≈qSBPT_{\text{fill}}(q)\approx\frac{qS}{B_P} and Tcpu(m−q)≈(m−q)SBH−BPT_{\text{cpu}}(m-q)\approx\frac{(m-q)S}{B_H-B_P}. Balancing the branches gives qm−q≈BPBH−BP\frac{q}{m-q}\approx\frac{B_P}{B_H-B_P} and therefore q∗≈mBPBHq^*\approx m\frac{B_P}{B_H}. The runtime rounds q∗q^* to an integer and retains at least one cache fill so that the cache continues warming. As BHB_H approaches BPB_P from above, q∗q^* approaches mm, reducing the policy to pure on-demand cache fill.

0

1

Updated 2026-09-12

Tags

Prep Sessions

Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.3 Adaptive Runtime Policies and Device Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Bandwidth-Adaptive Decode and the q* Execution Policy - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor