1Cademy - An engineer is deploying a large autoregressive model for a chatbot. They observe that as a conversation with a user gets longer, the models memory consumption increases steadily, eventually leading to performance issues. This is because the model stores key and value vectors for every token in the conversation history to speed up the generation of the next token. Based on this mechanism, what is the fundamental relationship between the length of the conversation history (in tokens) and the amount of memory required for this storage?

Learn Before

Space Complexity of the KV Cache

Multiple Choice

An engineer is deploying a large autoregressive model for a chatbot. They observe that as a conversation with a user gets longer, the model's memory consumption increases steadily, eventually leading to performance issues. This is because the model stores key and value vectors for every token in the conversation history to speed up the generation of the next token. Based on this mechanism, what is the fundamental relationship between the length of the conversation history (in tokens) and the amount of memory required for this storage?

Updated 2025-09-28

Contributors are:

Who are from:

Learn Before

Related