Plain versus Residual Baseline Architectures
The ImageNet plain baselines are primarily inspired by VGG networks. Their convolutional layers mostly use filters, with the same number of filters at a given feature-map size and twice as many filters whenever the feature-map size is halved. Downsampling uses stride-2 convolutions. Global average pooling is followed by a 1000-way fully connected layer and softmax. The 34-layer plain baseline requires 3.6 billion FLOPs, about 18% of VGG-19's 19.6 billion FLOPs.
Increasing the plain network from 18 to 34 layers raises validation top-1 error from 27.94% to 28.54%, illustrating the degradation problem. This optimization difficulty was not attributed to vanishing gradients: Batch Normalization was used, and forward signal variances and backward gradient norms remained healthy. Residual counterparts add shortcut connections across pairs of filters, using identity shortcuts and zero-padding when dimensions change. ResNet-34 achieves 25.03% top-1 error, compared with 27.88% for ResNet-18 and 28.54% for the 34-layer plain network. Its lower training error throughout optimization demonstrates that the residual formulation allowed the tested deeper network to benefit from added depth.
0
1
Tags
Prep Sessions
Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Ch.4 Residual Network Experiments and Applications - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
ImageNet Classification and Model Variations - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Learn After
In the plain baseline architecture inspired by VGG, what rule is applied to the filter count when the feature map size is halved?
In plain baseline architectures, the degradation problem observed when increasing network depth is caused by vanishing gradients.
How does the plain baseline network perform downsampling, and what layer does it terminate with immediately before the 1000-way fully-connected layer?
Compare the performance of plain networks and residual networks when scaling depth from 18 to 34 layers on ImageNet. In your response, describe the degradation problem observed in the plain baselines, how residual networks resolve this issue, and what training behavior demonstrates that residual networks effectively benefit from added depth.