Anaphora Resolution in Self-Attention Heads
In Transformer encoder self-attention, specialized attention heads participate directly in anaphora resolution by mapping pronouns and possessive determiners to their referents. In layer 5 of 6, isolated attention weights from words such as "its" show very sharp, focused connections to the relevant referent entities in the sequence.
0
1
Tags
Prep Sessions
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.1 Transformer Architecture Fundamentals - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Attention Visualizations and Linguistic Structure Resolution - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Related
Specialization of Self-Attention Heads for Linguistic Structure
Long-Distance Dependency Tracking in Self-Attention Heads
Anaphora Resolution in Self-Attention Heads
Match each aspect of self-attention analysis to its corresponding description.
Order the observations made when analyzing self-attention behavior across Transformer encoder heads, from initial visualization to structural interpretation.
According to findings from self-attention visualizations, what explains why the individual heads at layer 5 exhibit non-identical distributions?
Long-Distance Dependency Tracking in Self-Attention Heads
Anaphora Resolution in Self-Attention Heads
Learn After
In a 6-layer Transformer encoder, what characteristic pattern do isolated attention weights from words such as "its" display in layer 5?
In Transformer encoder self-attention, what specific linguistic process is performed by specialized attention heads that map pronouns and possessive determiners to their referents?