Sinusoidal versus Learned Positional Embeddings
Beyond fixed sinusoids, the architecture can alternatively utilize learned positional embeddings. In experimental variations on the English-to-German development set (newstest2013), replacing sinusoidal encodings with learned positional embeddings produced nearly identical performance, yielding a perplexity of 4.92 and a BLEU score of 25.7 compared to the base model's 4.92 perplexity and 25.8 BLEU score.
Despite the comparable empirical results, the sinusoidal version was selected for the final architecture because it provides an architectural advantage: fixed sinusoidal functions may allow the model to extrapolate to sequence lengths longer than any encountered during training.
0
1
Tags
Prep Sessions
Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Ch.1 Transformer Architecture and Components - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Sinusoidal and Learned Positional Encodings - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Learn After
When evaluating learned positional embeddings against the base sinusoidal model on the newstest2013 English-to-German development set, what was the empirical outcome?
Explain the sequence length extrapolation capability of fixed sinusoidal positional encodings in the Transformer architecture. In your response, define what this capability entails and explain why it served as the deciding factor in selecting sinusoidal encodings over learned positional embeddings.