1Cademy - An inference server is managing a batch of several short, ongoing requests that are in the process of generating output. A new request with a very long input sequence arrives. The systems scheduler immediately incorporates this new request into the active batch to begin processing it, aiming to keep the hardware as busy as possible. What is the most probable consequence for the initial short requests already in the batch?

Learn Before

Prefilling-Prioritized Strategy in Continuous Batching

Multiple Choice

An inference server is managing a batch of several short, ongoing requests that are in the process of generating output. A new request with a very long input sequence arrives. The system's scheduler immediately incorporates this new request into the active batch to begin processing it, aiming to keep the hardware as busy as possible. What is the most probable consequence for the initial short requests already in the batch?

Updated 2025-09-28

Contributors are:

Who are from:

Learn Before

Related