Discuss the significance of long-distance dependency tracking in a Transformer encoder. How does the ability of self-attention heads to link separated components support linguistic structure resolution across a sentence?
0
1
Tags
Prep Sessions
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.1 Transformer Architecture Fundamentals - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Attention Visualizations and Linguistic Structure Resolution - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Related
How do encoder self-attention heads resolve syntactic relationships when phrase components are separated by intervening words in a sentence?
True or False: In a 6-layer Transformer encoder, tracking long-distance dependencies between separated phrase components is performed exclusively by a single isolated attention head.
In the example illustrating long-distance dependency tracking across distant token positions, which phrase component is connected to the verb "making"?
Discuss the significance of long-distance dependency tracking in a Transformer encoder. How does the ability of self-attention heads to link separated components support linguistic structure resolution across a sentence?