Optimization Difficulties in Deep Plain Networks
Investigations of deep plain networks indicate that degradation is not attributable to vanishing or exploding signals. With batch normalization, forward-propagated signals retain nonzero variance and backward-propagated gradients retain healthy norms. Extending training to three times () the original number of iterations also fails to resolve the degradation. These observations motivate the conjecture that convergence in deep plain networks may be exponentially slow. They further suggest that standard optimizers have difficulty making stacks of nonlinear layers approximate identity mappings, whereas explicit identity shortcuts let the layers learn residual perturbations relative to the identity.
0
1
Tags
Prep Sessions
Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Ch.3 Deep Residual Network Architecture - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
The Degradation Problem in Deep Networks - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Learn After
What empirical observation demonstrates that the degradation problem in deep plain networks trained with batch normalization is not caused by vanishing or exploding signals?
Tripling the number of training iterations () effectively resolves the degradation problem in deep plain networks.
What specific form of convergence rate impedes training error reduction in deep plain networks suffering from the degradation problem?
Contrast the challenge standard optimizers face when an identity mapping is optimal in deep plain networks versus when using an explicit identity shortcut.