Comparison

Smaller ResNet Layer Responses and the Identity-Mapping Motivation

On CIFAR-10, the standard deviations of layer responses are generally smaller in ResNets than in corresponding plain networks. In a residual block, the learned branch represents F(x):=H(x)−x\mathcal{F}(\mathbf{x}) := \mathcal{H}(\mathbf{x}) - \mathbf{x}. If the desired mapping H\mathcal{H} is close to the identity mapping, then the residual mapping F\mathcal{F} is close to zero: the identity shortcut carries x\mathbf{x} directly while the residual branch learns a perturbation. The smaller measured responses are therefore consistent with, but do not by themselves prove, the identity-mapping motivation for residual learning.

0

1

Updated 2026-09-12

Tags

Prep Sessions

Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor

Ch.4 Residual Network Experiments and Applications - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor

Layer Response Analysis on CIFAR-10 - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor