Decode Throughput Stability in Agentic Serving
In multi-turn agentic workloads, inference sessions continuously expand their context length and invoke repeated tool interactions. Under these conditions, traditional Mixture-of-Experts (MoE) serving engines suffer significant decode throughput degradation compared to single-turn benchmarks (for example, losing more than 30% of generation throughput between single-turn reasoning and multi-turn coding). By contrast, serving systems utilizing dynamic LRU expert caching and bandwidth-adaptive miss partitioning maintain stable decode throughput, staying within 12% of single-turn performance across complex agent tasks. Because static offloading baselines deteriorate as conversations lengthen, single-stream, single-turn evaluations substantially overestimate baseline performance in real agentic deployments.
0
1
Tags
Prep Sessions
Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Ch.4 Platform Adaptation and Performance Evaluation - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Agentic Workload Serving and Cross-Hardware Performance - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor
Learn After
Why do single-stream, single-turn benchmark evaluations substantially overestimate baseline serving performance in agentic deployments?
In traditional Mixture-of-Experts (MoE) serving engines, generation throughput drops by more than what percentage when transitioning from single-turn reasoning to multi-turn coding tasks?
Analyze why static offloading baselines deteriorate in multi-turn agentic serving environments, and explain how advanced expert management techniques achieve stable decode throughput across extended sessions.