Self-Attention vs. Recurrent Layers: Complexity and Parallelization
For sequence length and representation dimension , a self-attention layer requires computation and sequential operations. A standard recurrent layer requires computation and sequential operations because each hidden state depends on . Thus, when , as is typical for word-piece or byte-pair sequences in the supplied course context, self-attention has lower asymptotic per-layer complexity and permits parallel computation across sequence positions. For long sequences, restricting each position to an attention neighborhood of size reduces the per-layer complexity to while retaining sequential operations.
0
1
Tags
Prep Sessions
Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Ch.2 Transformer Training and Evaluation - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Complexity and Path Lengths in Self-Attention - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Learn After
In natural language processing tasks using word-piece or byte-pair representations, why are self-attention layers typically computationally faster than recurrent layers?
In self-attention layers, computing representations requires waiting for earlier sequence positions to complete before later positions can be processed.
Why do standard recurrent layers require sequential operations to process a sequence of length ?
Explain the motivation, mechanism, and resulting computational profile (both per-layer complexity and sequential operations) of restricted self-attention for long sequences.