Concept icon
Concept

Optimization Difficulties in Deep Plain Networks

Investigations of deep plain networks indicate that degradation is not attributable to vanishing or exploding signals. With batch normalization, forward-propagated signals retain nonzero variance and backward-propagated gradients retain healthy norms. Extending training to three times (3×3\times) the original number of iterations also fails to resolve the degradation. These observations motivate the conjecture that convergence in deep plain networks may be exponentially slow. They further suggest that standard optimizers have difficulty making stacks of nonlinear layers approximate identity mappings, whereas explicit identity shortcuts let the layers learn residual perturbations relative to the identity.

0

1

Concept icon
Updated 2026-09-19

Tags

Prep Sessions

Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor

Ch.3 Deep Residual Network Architecture - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor

The Degradation Problem in Deep Networks - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor