Concept icon
Concept

Decode Throughput Stability in Agentic Serving

In multi-turn agentic workloads, inference sessions continuously expand their context length and invoke repeated tool interactions. Under these conditions, traditional Mixture-of-Experts (MoE) serving engines suffer significant decode throughput degradation compared to single-turn benchmarks (for example, losing more than 30% of generation throughput between single-turn reasoning and multi-turn coding). By contrast, serving systems utilizing dynamic LRU expert caching and bandwidth-adaptive miss partitioning maintain stable decode throughput, staying within 12% of single-turn performance across complex agent tasks. Because static offloading baselines deteriorate as conversations lengthen, single-stream, single-turn evaluations substantially overestimate baseline performance in real agentic deployments.

0

1

Concept icon
Updated 2026-09-07

Tags

Prep Sessions

Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Ch.4 Platform Adaptation and Performance Evaluation - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor

Agentic Workload Serving and Cross-Hardware Performance - Edge-Native Mixture-of-Experts Serving with FreeToken @ University of Michigan - Ann Arbor