Learn Before
In the described training configuration, approximately how many tokens did each training batch contain?
0
1
Tags
Prep Sessions
Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Ch.2 Transformer Training and Evaluation - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Transformer Training, Regularization, and Optimization - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Related
Importance of Training Data Visualization
In the described training configuration, approximately how many tokens did each training batch contain?
The standard WMT 2014 English-German training dataset contained approximately 36 million sentence pairs.
What optimizer was specified for training, and what values were used for its beta 1 (B1), beta 2 (B2), and epsilon (e) hyperparameters?
Detail the two datasets used in training, including their sizes, how the English-German data was encoded, and how sentence pairs were batched.