Learn Before
Residual Formulation and Shortcut Connections - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Identity versus Projection Shortcuts - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
ImageNet Classification and Model Variations - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Dimension Matching and Projection Shortcuts
The element-wise addition in a residual block requires that the dimensions of the input and the residual output match. When changing channel dimensions or downsampling feature map sizes across stages (performed with a stride of 2), a dimension-matching operation must be applied to the shortcut.
To match dimensions, a linear projection is applied along the shortcut connection: Three shortcut strategies are evaluated for handling dimension changes:
- Option A: Zero-padding shortcuts are used to increase dimensions, keeping all shortcuts parameter-free.
- Option B: Projection shortcuts (implemented via convolutions) are used only when increasing dimensions, while all other shortcuts remain parameter-free identity mappings.
- Option C: All shortcuts in the network are linear projections.
While all three options substantially outperform plain baselines, Option B achieves slightly better accuracy than Option A because zero-padded entries do not participate in residual learning. Option C yields only a marginal improvement over Option B at the expense of adding many extra parameters. Because projection shortcuts are not essential for solving the degradation problem, identity shortcuts are preferred to minimize model size and computational complexity.
0
1
Tags
Prep Sessions
Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Ch.3 Deep Residual Network Architecture - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Residual Formulation and Shortcut Connections - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Identity versus Projection Shortcuts - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Ch.4 Residual Network Experiments and Applications - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
ImageNet Classification and Model Variations - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Related
Dimension Matching and Projection Shortcuts
Bottleneck Residual Blocks
Residual Mapping
Shortcut’s technique for identity mapping
Dimension Matching and Projection Shortcuts
Empirical Performance of Identity vs. Projection Shortcuts
Bottleneck Residual Blocks
Plain versus Residual Baseline Architectures
Dimension Matching and Projection Shortcuts
Deep ResNet Scaling with Bottleneck Blocks
ImageNet Classification Results and Benchmarks
Learn After
Match each shortcut strategy for handling dimension changes to its corresponding implementation rule.
Order the operations performed to compute the output of a residual block with a projection shortcut when changing dimensions.
Explain why Option B achieves higher accuracy than Option A, and state which type of shortcut should be favored whenever dimensions match to minimize model size and computational complexity.
Match each mathematical or structural element of dimension matching to its specific role in residual networks.
Order the shortcut strategies by the total number of parameters they add to the network's shortcut connections, from fewest (least parameters) to most (greatest parameters).
Based on the evaluated shortcut strategies, evaluate whether switching from Option B to Option C is justified in terms of accuracy and model complexity, and explain why identity shortcuts are generally favored.
Match each shortcut evaluation strategy to its corresponding performance trade-off and empirical finding in residual networks.
When downsampling feature map sizes across stages in a residual network, the operation is performed with a stride of ___ .
Explain why the shortcut input cannot be combined directly with the residual output during this stage transition, and describe how applying a linear projection resolves this issue.