Concept icon
Concept

Vector Products per Self-Attention Step

During a single step of standard autoregressive generation, attending a position i′i' to all previous context positions requires exactly 2i′{}2 i' vector products. This total is comprised of i′i' products needed for the query-key dot product (qi′KT\mathbf{q}_{i'} \mathbf{K}^{\mathrm{T}}), plus an additional i′i' products to multiply the Softmax-normalized attention scores with the value matrix (Softmax(qi′KTd)V\mathrm{Softmax}(\frac{\mathbf{q}_{i'} \mathbf{K}^{\mathrm{T}}}{\sqrt{d}}) \mathbf{V}).

0

1

Concept icon
Updated 2026-05-03

Contributors are:

Who are from:

Tags

Foundations of Large Language Models

Ch.5 Inference - Foundations of Large Language Models

Foundations of Large Language Models Course

Computing Sciences