Concept icon
Concept

Cross-Hardware MoE Serving across Memory and Interconnect Tiers

The practical performance of hybrid MoE execution across consumer systems depends directly on the relative balance between host memory bandwidth (BHB_H) and PCIe transfer bandwidth (BPB_P). On dual-channel consumer desktop platforms, CPU-side expert execution encounters a memory bandwidth ceiling (approximately 5054 GB/s50\text{--}54\text{ GB/s}), causing CPU-bound baselines to lose 20% or more of their generation throughput compared to multi-channel server configurations. By contrast, bandwidth-adaptive serving dynamically routes miss handling toward PCIe-driven GPU cache fills when interconnect bandwidth is high relative to host RAM, allowing consumer desktops to preserve up to 96% of multi-channel server throughput. Across heterogeneous form factors—from an 8 GB laptop GPU over PCIe 4.0 x8 to a workstation GPU running 700B+ models—adapting miss distribution to measured bandwidth ratios delivers consistent acceleration over static offloading schemes.

0

1

Concept icon
Updated 2026-09-07

Tags

Prep Sessions

Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.4 Platform Adaptation and Performance Evaluation - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Agentic Workload Serving and Cross-Hardware Performance - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor