Transformer Adam Settings and Warmup–Inverse-Square-Root Learning Rate Formula
Transformer training uses the Adam optimizer with , , and . At step , the learning rate is where is the model dimension and is the number of warmup steps. The rate increases linearly during warmup and then decays in proportion to . The reported experiments use .
0
1
Tags
Prep Sessions
Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Ch.2 Transformer Training and Evaluation - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Transformer Training, Regularization, and Optimization - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Related
Training Data
Regularization Techniques in the Transformer Model
Transformer Adam Settings and Warmup–Inverse-Square-Root Learning Rate Formula
Effect of Learning Rate Scheduling on Overfitting
Polynomial Learning Rate Decay
Piecewise Constant Learning Rate Schedule
Cosine Learning Rate Schedule
Optimizer Warmup
Factor Learning Rate Scheduler
Explicit Learning Rate Adjustment Implementation
Learning Rate Scheduler Toy Problem
Square Root Learning Rate Scheduler
Transformer Adam Settings and Warmup–Inverse-Square-Root Learning Rate Formula
Learn After
Match each optimizer hyperparameter or schedule parameter to its specific value used in the reported Transformer experiments.
During the initial phase defined by warmup_steps, the Transformer schedule modulates the learning rate such that it increases ___ with respect to the step number.
Order the trajectory of the learning rate across training steps from earliest to latest.
Explain why Model B has a lower learning rate than Model A throughout training according to the scheduling formula, and calculate the exact ratio of Model B's learning rate to Model A's learning rate (lrate_B / lrate_A).