Essay

Analyze the effect of varying the number of attention heads (hh) on Transformer performance based on the ablation experiments on the English-to-German development set (newstest2013). Detail how single-head attention, the baseline 8-head configuration, and an excessively high head count compare.

0

1

Updated 2026-09-07

Tags

Prep Sessions

Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor

Ch.2 Transformer Training and Evaluation - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor

Machine Translation and Constituency Parsing Evaluation - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor