Residual Responses and Identity Preconditioning
Empirical analysis on CIFAR-10 reveals that the standard deviations of layer responses in ResNets are generally smaller than those in corresponding plain networks.
This behavior supports the foundational premise of deep residual learning. Formulating the network layers to learn residual mappings assumes that optimal underlying mappings are typically closer to identity mappings than to zero mappings. Because identity shortcuts preserve information directly, the solver only needs to learn small perturbations with reference to the identity mapping, resulting in residual function responses that stay close to zero.
0
1
Tags
Prep Sessions
Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Ch.4 Residual Network Experiments and Applications - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Layer Response Analysis on CIFAR-10 - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Learn After
Match each mathematical notation or structural component of deep residual learning to its functional definition.
According to the foundational premise of residual learning, why do the residual function responses stay close to zero and produce smaller response standard deviations compared to the plain network?