Path Lengths and Long-Range Dependency Learning
A model's ability to learn long-range dependencies is affected by the maximum path length that signals must traverse between arbitrary input and output positions. Shorter paths reduce the number of transformations traversed and can make long-range dependencies easier to learn. Full self-attention directly connects every position to every other position, giving a maximum path length of . Recurrent architectures process positions sequentially and have a maximum path length of . For convolutional layers with kernel width , connecting all positions requires stacked layers with contiguous kernels or layers with dilated convolutions. Restricted self-attention over neighborhoods of size has a maximum path length of .
0
1
Tags
Prep Sessions
Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Ch.2 Transformer Training and Evaluation - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Complexity and Path Lengths in Self-Attention - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Learn After
Match each layer architecture to its corresponding maximum path length between arbitrary input and output positions.
Connecting all pairs of input and output positions using a stack of convolutional layers with contiguous kernels of width k requires a maximum path length of ___.
Order these layer architectures from shortest to longest maximum path length between arbitrary input and output positions (assuming kernel width 1 < k < n):
Based on path length analysis, explain how adopting restricted self-attention changes the maximum path length compared to full self-attention, and describe the resulting effect on gradient flow and long-range dependency learning.