1Cademy - Calculating a Linear Attention Output Vector

Learn Before

Linear Attention Output Calculation
Linear Causal Attention Formula

Short Answer

Calculating a Linear Attention Output Vector

In a linear attention mechanism, at a specific timestep i, you are given the following:

Transformed query vector: q'_i = [2, 1]
Accumulated key-value state: μ_i = [[10, 5], [4, 8]]
Accumulated key state: ν_i = [6, 3]

Using the formula Output = (q'_i * μ_i) / (q'_i * ν_i), calculate the final output vector. Provide the calculated values for the numerator and the denominator before giving the final result.

Updated 2026-04-22

Contributors are:

Who are from:

Tags

Ch.2 Generative Models - Foundations of Large Language Models

Foundations of Large Language Models

Foundations of Large Language Models Course

Computing Sciences

Application in Bloom's Taxonomy

Cognitive Psychology

Psychology

Social Science

Empirical Science

Science

In the formula for calculating a linear attention output, Output = (q'_i * μ_i) / (q'_i * ν_i), where q'_i is the transformed query, μ_i is the accumulated key-value state, and ν_i is the accumulated key state, what is the primary role of the denominator term q'_i * ν_i?
Calculating a Linear Attention Output Vector
In a memory-efficient attention mechanism, the output for a token at position i is calculated using the formula: Output = (q'_i * μ_i) / (q'_i * ν_i). In this formula, q'_i is the token's processed query, while μ_i and ν_i are aggregations of historical information from all tokens up to and including position i. Specifically, μ_i aggregates past key-value products, and ν_i aggregates past keys. What is the primary function of the denominator, q'_i * ν_i?
Efficiency of Aggregated State in Attention
Evaluating a Modification to the Linear Attention Formula
In the formula for calculating a linear attention output, Output = (q'_i * μ_i) / (q'_i * ν_i), where q'_i is the transformed query, μ_i is the accumulated key-value state, and ν_i is the accumulated key state, what is the primary role of the denominator term q'_i * ν_i?
Calculating a Linear Attention Output Vector
Recurrent Computation of $\mu_i$ and $\nu_i$ in Linear Attention

Learn Before

Related