Long-Distance Dependency Tracking in Self-Attention Heads
Encoder self-attention heads can track long-distance syntactic dependencies across a sentence. As demonstrated in layer 5 of a 6-layer Transformer encoder, multiple attention heads attend across intervening words to link separated components of a phrase, such as connecting the verb "making" to "more difficult" across distant token positions.
0
1
Tags
Prep Sessions
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.1 Transformer Architecture Fundamentals - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Attention Visualizations and Linguistic Structure Resolution - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Related
Specialization of Self-Attention Heads for Linguistic Structure
Long-Distance Dependency Tracking in Self-Attention Heads
Anaphora Resolution in Self-Attention Heads
Match each aspect of self-attention analysis to its corresponding description.
Order the observations made when analyzing self-attention behavior across Transformer encoder heads, from initial visualization to structural interpretation.
According to findings from self-attention visualizations, what explains why the individual heads at layer 5 exhibit non-identical distributions?
Long-Distance Dependency Tracking in Self-Attention Heads
Anaphora Resolution in Self-Attention Heads
Learn After
How do encoder self-attention heads resolve syntactic relationships when phrase components are separated by intervening words in a sentence?
True or False: In a 6-layer Transformer encoder, tracking long-distance dependencies between separated phrase components is performed exclusively by a single isolated attention head.
In the example illustrating long-distance dependency tracking across distant token positions, which phrase component is connected to the verb "making"?
Discuss the significance of long-distance dependency tracking in a Transformer encoder. How does the ability of self-attention heads to link separated components support linguistic structure resolution across a sentence?