Comparison

Path Lengths and Long-Range Dependency Learning

A model's ability to learn long-range dependencies is affected by the maximum path length that signals must traverse between arbitrary input and output positions. Shorter paths reduce the number of transformations traversed and can make long-range dependencies easier to learn. Full self-attention directly connects every position to every other position, giving a maximum path length of O(1)O(1). Recurrent architectures process positions sequentially and have a maximum path length of O(n)O(n). For convolutional layers with kernel width k<nk<n, connecting all positions requires O(n/k)O(n/k) stacked layers with contiguous kernels or O(log⁡kn)O(\log_k n) layers with dilated convolutions. Restricted self-attention over neighborhoods of size rr has a maximum path length of O(n/r)O(n/r).

0

1

Updated 2026-09-19

Tags

Prep Sessions

Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor

Ch.2 Transformer Training and Evaluation - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor

Complexity and Path Lengths in Self-Attention - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor