Autoregressive Dependency Guarantee in Transformer Decoders
To preserve autoregression in the Transformer decoder, the prediction generated for position must depend exclusively on known outputs at positions strictly less than . This constraint is enforced by two synchronized mechanisms: modifying the decoder's self-attention sub-layer to mask out subsequent positions and shifting the output embeddings to the right by one position.
0
1
Tags
Prep Sessions
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.1 Transformer Architecture Fundamentals - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Decoder Masking and Autoregressive Prediction - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Related
Masked Self-Attention in Transformer Decoders
Autoregressive Dependency Guarantee in Transformer Decoders
In a Transformer decoder, masked self-attention is used to ensure that the prediction for a token at a given position can only depend on previous tokens. This is achieved by modifying the attention score matrix before the softmax function is applied. For a sequence of tokens, which of the following correctly describes the structure of the attention score matrix after this causal mask has been applied?
In masked self-attention, which key positions is the query of a given token permitted to interact with?
In standard self-attention, each position is restricted from attending to subsequent tokens in the sequence.
Under masked self-attention, which key positions is a query for a given token permitted to interact with?
What core capability does masked self-attention enable in the Transformer decoder?
Order the operations performed during masked self-attention from first to last:
Autoregressive Dependency Guarantee in Transformer Decoders
Learn After
To preserve autoregression in the Transformer decoder, what dependency constraint must be satisfied by the prediction generated for position ?
What positional adjustment must be applied to the output embeddings to help enforce the autoregressive constraint in the Transformer decoder?
Explain how the Transformer decoder prevents future position information from influencing predictions, detailing the necessary mechanisms and the operational condition they uphold.