Training Data
- Training data included the standard WMT 2014 English-German dataset consisting of about 4.5 million sentence pairs. Sentences were encoded using byte-pair encoding.
- For English-French, we used the significantly larger WMT 2014 English-French dataset consisting of 36M. Sentence pairs were batched together by approximate sequence length.
- Each training batch contained a set of sentence pairs containing approximately 25000 source tokens and 25000 target tokens.
- Optimizer – Adam B1 = 0:9, B2 = 0:98 and e = 10^-9
0
1
Contributors are:
Who are from:
Tags
Data Science
Prep Sessions
Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Ch.2 Transformer Training and Evaluation - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Transformer Training, Regularization, and Optimization - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Related
What problem is this paper trying to solve?
Transformer model
Training Data
Advantages and Performance of the Transformer Model
Regularization Techniques in the Transformer Model
Attention Functions Used in the Transformer Model
Optimizer and Learning Rate Schedule
Training Data
Regularization Techniques in the Transformer Model
Learn After
Importance of Training Data Visualization
In the described training configuration, approximately how many tokens did each training batch contain?
The standard WMT 2014 English-German training dataset contained approximately 36 million sentence pairs.
What optimizer was specified for training, and what values were used for its beta 1 (B1), beta 2 (B2), and epsilon (e) hyperparameters?
Detail the two datasets used in training, including their sizes, how the English-German data was encoded, and how sentence pairs were batched.