Concept icon
Concept

Cold-Cache Serving Without GPU Warmup

Traditional serving runtimes require a dedicated warmup phase prior to processing user requests to prime GPU execution caches and allocators, incurring substantial startup delay on consumer edge devices. By contrast, cold-cache serving eliminates pre-warming entirely by construction. The engine begins serving the first incoming request immediately with an unpopulated, cold GPU expert cache. Initial expert requests naturally trigger cache misses that are processed directly by the standard decode miss path through concurrent PCIe cache fills and in-place CPU execution. As generation proceeds, the expert cache warms organically during ordinary inference without requiring synthetic warmup passes.

0

1

Concept icon
Updated 2026-09-07

Tags

Prep Sessions

Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.3 Adaptive Runtime Policies and Device Execution - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Elastic Runtime Cache Reconfiguration and Fast Bootstrap - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor