Formula

Transformer Adam Settings and Warmup–Inverse-Square-Root Learning Rate Formula

Transformer training uses the Adam optimizer with β1=0.9\beta_1=0.9, β2=0.98\beta_2=0.98, and ϵ=10−9\epsilon=10^{-9}. At step ss, the learning rate is lrate⁡(s)=dmodel−0.5min⁡(s−0.5,;s,w−1.5),\operatorname{lrate}(s)=d_{\text{model}}^{-0.5}\min\left(s^{-0.5},;s,w^{-1.5}\right), where dmodeld_{\text{model}} is the model dimension and ww is the number of warmup steps. The rate increases linearly during warmup and then decays in proportion to s−0.5s^{-0.5}. The reported experiments use w=4000w=4000.

0

1

Updated 2026-09-19

Tags

Prep Sessions

Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor

Ch.2 Transformer Training and Evaluation - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor

Transformer Training, Regularization, and Optimization - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor