Learn Before
In a 6-layer Transformer encoder, what characteristic pattern do isolated attention weights from words such as "its" display in layer 5?
0
1
Tags
Prep Sessions
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.1 Transformer Architecture Fundamentals - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Attention Visualizations and Linguistic Structure Resolution - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Related
In a 6-layer Transformer encoder, what characteristic pattern do isolated attention weights from words such as "its" display in layer 5?
In Transformer encoder self-attention, what specific linguistic process is performed by specialized attention heads that map pronouns and possessive determiners to their referents?