Comparison

Layer-Specific Parameters and Convolutional Interpretation of Transformer FFNs

A Transformer position-wise feed-forward network applies the same two-layer transformation independently at every token position within a layer, while its parameters W1,b1,W2,b2W_1, b_1, W_2, b_2 differ across layers. This position-wise transformation can be interpreted as two consecutive convolutions with kernel size 11. Combining self-attention with a point-wise feed-forward layer has the same computational complexity as a separable convolution whose kernel size kk equals the sequence length nn. In the encoder and decoder, the feed-forward sub-layer is wrapped by a residual connection and layer normalization, producing operatorname{LayerNorm}(x + operatorname{Sublayer}(x)).

0

1

Updated 2026-09-12

Tags

Prep Sessions

Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor

Ch.1 Transformer Architecture and Components - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor

Position-Wise Feed-Forward Networks - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor

Related