Concept icon
Concept

Residual Responses and Identity Preconditioning

Empirical analysis on CIFAR-10 reveals that the standard deviations of layer responses in ResNets are generally smaller than those in corresponding plain networks.

This behavior supports the foundational premise of deep residual learning. Formulating the network layers to learn residual mappings F(x):=H(x)x\mathcal{F}(\mathbf{x}) := \mathcal{H}(\mathbf{x}) - \mathbf{x} assumes that optimal underlying mappings are typically closer to identity mappings than to zero mappings. Because identity shortcuts preserve information directly, the solver only needs to learn small perturbations with reference to the identity mapping, resulting in residual function responses that stay close to zero.

0

1

Concept icon
Updated 2026-09-07

Tags

Prep Sessions

Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor

Ch.4 Residual Network Experiments and Applications - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor

Layer Response Analysis on CIFAR-10 - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor