Vector Mapping Definition of Attention
An attention function is defined as a mapping from a query and a set of key-value pairs to an output, where the query, keys, values, and resulting output are all represented as vectors.
0
1
Tags
Prep Sessions
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.1 Transformer Architecture Fundamentals - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Self-Attention and Query-Key-Value Mechanisms - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Related
Self-Attention
Vector Mapping Definition of Attention
Query, Key, and Value in Attention Mechanisms
Query (Attention)
Key (Attention)
Value (Attention)
State Function from Previous Outputs
Value Weight Matrix Formula
Set of Sequential Key-Value Pairs
Query Vector
Key Vector
Value Vector
Implicit Relative Position Modeling in Self-Attention with RoPE
Imagine a system translating the sentence 'The quick brown fox jumps'. When the system is generating the output word corresponding to 'jumps', it needs to determine which words in the input sentence are most relevant. To do this, a vector representing the current translation context (i.e., 'what information do I need to produce the next word?') is compared against a set of searchable 'label' vectors, one for each word in the input sentence. This comparison generates a relevance score for each in
In a system designed to answer questions based on a provided document, the model first creates a representation of the user's question. It then compares this representation against a set of searchable representations, one for each sentence in the document, to determine relevance scores. Finally, it constructs an answer by creating a weighted combination of the informational content from each sentence, using the relevance scores as weights. Which option correctly assigns the roles of Query, Key,
Context Window of Key Vectors Notation
Key-Value Cache
In a computational mechanism designed to selectively focus on different parts of an input sequence, information is represented by three distinct types of vectors that interact to produce a context-aware output. Match each vector type to its specific role in this process.
Masked QKV Attention Formula
Nadaraya-Watson Estimator
In an attention mechanism, weights computed by comparing a query with a set of keys are applied to the corresponding ___ to produce an output.
Based on how attention components interact, identify which value will dominate the resulting output and explain how the query-key comparisons lead to this outcome.
Vector Mapping Definition of Attention
Learn After
In the formal definition of attention, how are the keys and values structured when provided as inputs to the mapping function?
True or False: In the formal definition of attention, the attention function maps from multiple queries simultaneously to produce a single output.
In the formal definition of an attention function, what mathematical representation is uniformly shared by the query, the keys, the values, and the resulting output?