Learn Before
Transformer Decoder
Decoder Masking and Autoregressive Prediction - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Scaled Dot-Product Attention - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Masked Self-Attention in Transformer Decoders
Masked self-attention is a crucial component of the Transformer decoder, enabling autoregressive text generation. Unlike standard self-attention, it restricts each position from attending to subsequent, or 'future,' positions in the sequence. This is implemented by applying a mask to the attention scores before the softmax function, effectively zeroing out the weights for future tokens. Consequently, the query for a given token can only interact with keys from its own position and all preceding positions, ensuring that the prediction for the current step depends only on the known past.
0
1
Contributors are:
Who are from:
References
The Illustrated Transformer
Attention Is All You Need
Reference of Foundations of Large Language Models Course
Reference of Foundations of Large Language Models Course
Reference of Foundations of Large Language Models Course
Reference of Foundations of Large Language Models Course
Dive into Deep Learning
attention-first3-last3.pdf
Tags
Data Science
Ch.5 Inference - Foundations of Large Language Models
Foundations of Large Language Models
Foundations of Large Language Models Course
Computing Sciences
Ch.2 Generative Models - Foundations of Large Language Models
D2L
Dive into Deep Learning @ D2L
Prep Sessions
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.1 Transformer Architecture Fundamentals - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Decoder Masking and Autoregressive Prediction - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Ch.1 Transformer Architecture and Components - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Scaled Dot-Product Attention - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Related
Core Components of a Transformer Decoding Network
Masked Self-Attention in Transformer Decoders
A standard Transformer decoder block contains two distinct attention sub-layers. Which statement accurately differentiates the roles and data sources for these two sub-layers?
Within a single decoder block of a standard Transformer architecture, information is processed through three main computational sub-layers. Arrange these sub-layers in the correct operational sequence.
In the Transformer decoder's encoder-decoder attention layer, what is the decoder query compared with?
What is the defining rule of the masked self-attention layer in the Transformer decoder?
Autoregressive Dependency Guarantee in Transformer Decoders
Masked Self-Attention in Transformer Decoders
Scaled Dot-Product Attention
Variance Control in Dot Product Attention
Masked Self-Attention in Transformer Decoders
Learn After
In a Transformer decoder, masked self-attention is used to ensure that the prediction for a token at a given position can only depend on previous tokens. This is achieved by modifying the attention score matrix before the softmax function is applied. For a sequence of tokens, which of the following correctly describes the structure of the attention score matrix after this causal mask has been applied?
In masked self-attention, which key positions is the query of a given token permitted to interact with?
Autoregressive Dependency Guarantee in Transformer Decoders
In standard self-attention, each position is restricted from attending to subsequent tokens in the sequence.
Under masked self-attention, which key positions is a query for a given token permitted to interact with?
What core capability does masked self-attention enable in the Transformer decoder?
Order the operations performed during masked self-attention from first to last: