Learn Before
Short Answer

What operation is applied to the individual parallel head outputs (head1,,headh\text{head}_1, \dots, \text{head}_h) immediately before multiplying them by the parameter matrix WOW^O?

0

1

Updated 2026-09-07

Tags

Prep Sessions

Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor

Ch.1 Transformer Architecture and Components - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor

Multi-Head Attention - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor