Analyze how the hardware differences between an LPDDR5 laptop and a high-speed PCIe desktop affect the strategy for serving MoE decode cache misses. Ground your analysis in how host memory bandwidth and PCIe transfer bandwidth interact across these platforms.
0
1
Tags
Prep Sessions
Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Ch.1 Edge Serving Bottlenecks and Dynamics - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Decode Cache Misses and Host Bandwidth Bottlenecks - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Ch.3 Adaptive Runtime Policies and Device Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Bandwidth-Adaptive Decode and the q* Execution Policy - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Related
During MoE decoding, under what specific circumstance does the system need to decide between transferring expert weights across PCIe and computing directly on the CPU?
When an active expert is computed entirely on the CPU during a decode cache miss, the system still receives the benefit of future cache hits from a GPU VRAM cache fill.
Under what condition does exclusively transferring expert weights across PCIe to the GPU leave host DRAM bandwidth and CPU cores idle?
Analyze how the hardware differences between an LPDDR5 laptop and a high-speed PCIe desktop affect the strategy for serving MoE decode cache misses. Ground your analysis in how host memory bandwidth and PCIe transfer bandwidth interact across these platforms.
Why can the optimal strategy for dividing MoE decode cache miss handling between CPU compute and PCIe transfers not be statically predetermined across all systems?
When an active expert is computed entirely on the CPU during an MoE decode cache miss, which hardware interconnect is left idle?
Cross-Hardware MoE Serving across Memory and Interconnect Tiers