Learn Before
Order the steps taken to compute a stabilized attention score between a query vector and a key vector.
0
1
Tags
Prep Sessions
Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Ch.1 Transformer Architecture and Components - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Scaled Dot-Product Attention - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Related
Match each statistical property of dot product attention to its corresponding value, assuming vector elements are i.i.d. with a mean of 0 and a variance of 1.
Order the steps taken to compute a stabilized attention score between a query vector and a key vector.
Explain why unscaled dot products cause vanishing gradients in this model, and calculate the exact numerical scaling factor required to maintain a variance of 1.