Concept icon
Concept

Anaphora Resolution in Self-Attention Heads

In Transformer encoder self-attention, specialized attention heads participate directly in anaphora resolution by mapping pronouns and possessive determiners to their referents. In layer 5 of 6, isolated attention weights from words such as "its" show very sharp, focused connections to the relevant referent entities in the sequence.

0

1

Concept icon
Updated 2026-09-07

Tags

Prep Sessions

Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor

Ch.1 Transformer Architecture Fundamentals - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor

Attention Visualizations and Linguistic Structure Resolution - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor