Activity (Process)

Fused Multi-Bank Expert Transfer via Device-Resident Work Lists

To transfer missing expert parameters across PCIe without host synchronization overhead, expert weights are structured so that all parameter banks share identical logical expert-to-slot mappings. Once victim slots are selected on the GPU, a copy work list is generated in device memory containing source and destination indices. This single index list is launched across all weight banks in a single, fused memory transfer kernel of fixed shape. Any unused slots within the fixed-dimension work buffer are masked out using a device-resident valid count, minimizing kernel launch overhead, sustaining maximum PCIe bus saturation, and keeping execution entirely device-resident.

0

1

Updated 2026-09-07

Tags

Prep Sessions

Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.3 Adaptive Runtime Policies and Device Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

CUDA-Graph-Compatible Device-Side Cache Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.4 Platform Adaptation and Performance Evaluation - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Expert Storage Formats and Platform Adaptation - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor