Regularization Techniques in the Transformer Model
The Transformer model employs two primary regularization techniques during training:
-
Residual Dropout: Applied to the output of each sub-layer before it is added to the sub-layer input and normalized. Dropout is also applied to the sums of the embeddings and the positional encodings in both the encoder and decoder stacks. For the base model, a dropout rate of is used.
-
Label Smoothing: Employed during training. While this hurts perplexity, as the model learns to be more unsure, it improves accuracy and BLEU (Bilingual Evaluation Understudy) scores.

0
1
Contributors are:
Who are from:
Tags
Data Science
Prep Sessions
Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Ch.2 Transformer Training and Evaluation - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Transformer Training, Regularization, and Optimization - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Related
What problem is this paper trying to solve?
Transformer model
Training Data
Advantages and Performance of the Transformer Model
Regularization Techniques in the Transformer Model
Attention Functions Used in the Transformer Model
Optimizer and Learning Rate Schedule
Training Data
Regularization Techniques in the Transformer Model
Learn After
Where is residual dropout applied within each sub-layer of the Transformer architecture?
In the Transformer architecture, dropout is applied to the sums of the embeddings and positional encodings in both the encoder and decoder stacks.
What is the dropout rate () utilized in the base Transformer model?
Describe the trade-offs of applying label smoothing during Transformer training. Discuss how label smoothing alters the model's behavior, its impact on perplexity, and its effect on final task metrics like BLEU score and accuracy.