Transformer Positional Encoding: Sinusoidal versus Learned Embeddings
The Transformer can use either fixed sinusoidal positional encodings or learned positional embeddings. On the newstest2013 English-to-German development set, the learned variant nearly matched the base sinusoidal model: both had a perplexity of 4.92, while BLEU was 25.7 for learned embeddings and 25.8 for the base model. The sinusoidal variant was retained because its fixed functions may extrapolate to sequence lengths beyond those encountered during training.
0
1
Tags
Prep Sessions
Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Ch.1 Transformer Architecture and Components - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Sinusoidal and Learned Positional Encodings - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Related
Positional Encoding
Sinusoidal Positional Encoding
Transformer Positional Encoding: Sinusoidal versus Learned Embeddings
A development team is building a language model that will be trained on documents with a maximum length of 512 tokens. However, a critical requirement for the final application is that the model must effectively process documents that are occasionally up to 4000 tokens long. The team chooses to use a position representation method based on a combination of sine and cosine functions of different frequencies. Which of the following statements most accurately evaluates this choice?
Analyzing the Trade-offs of Sinusoidal Positional Encoding
Match each mathematical component of the sinusoidal positional encoding scheme with its description.
Which two mathematical functions serve as the foundation for the fixed positional encoding scheme described?
Transformer Positional Encoding: Sinusoidal versus Learned Embeddings
Generalization Issues of Learnable Positional Embeddings
A language model is trained exclusively on text sequences with a maximum length of 512 tokens. This model uses a method where a unique vector is learned for each specific position in the sequence (e.g., a vector for position 1, a different vector for position 2, etc., up to position 512). After training is complete, the model is tasked with processing a new sequence that is 600 tokens long. What is the most direct and fundamental problem the model will encounter when processing the tokens from p
Analysis of Positional Vector Assignment
A language model architect is designing a system to process sequences with a maximum length of 1024 tokens. They opt for an approach where a unique vector is created for each position (1, 2, ..., 1024). These vectors are initialized randomly and are updated based on the training objective, just like the other parameters in the model. Which statement best analyzes a key characteristic of this specific method for encoding position?
Limitation of Independent Positional Embeddings
Transformer Positional Encoding: Sinusoidal versus Learned Embeddings
Learn After
When evaluating learned positional embeddings against the base sinusoidal model on the newstest2013 English-to-German development set, what was the empirical outcome?
Explain the sequence length extrapolation capability of fixed sinusoidal positional encodings in the Transformer architecture. In your response, define what this capability entails and explain why it served as the deciding factor in selecting sinusoidal encodings over learned positional embeddings.