Learn Before
Residual Formulation and Shortcut Connections - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Identity versus Projection Shortcuts - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
ImageNet Classification and Model Variations - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Identity versus Projection Shortcuts - Overcoming Neural Network Degradation Through Residual Learning @ University of Michigan - Ann Arbor
Residual Mapping
Dimension Matching and Projection Shortcuts
Element-wise addition in a residual block requires the shortcut output and the residual output to have the same dimensions. When channel dimensions increase or spatial feature maps are downsampled across stages, the shortcut must be adjusted before addition:
Three shortcut strategies handle dimension changes:
- Option A — zero-padding shortcut: Retain a parameter-free identity shortcut and pad additional channel entries with zeros.
- Option B — projection only for dimension increases: Apply a convolutional projection when dimensions change, using stride when spatial downsampling is required; otherwise use an identity shortcut.
- Option C — projection for all shortcuts: Replace every shortcut with a convolutional projection, even when input and output dimensions already match.
0
1
Tags
Prep Sessions
Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Ch.3 Deep Residual Network Architecture - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Residual Formulation and Shortcut Connections - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Identity versus Projection Shortcuts - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Ch.4 Residual Network Experiments and Applications - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
ImageNet Classification and Model Variations - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Overcoming Neural Network Degradation Through Residual Learning @ University of Michigan - Ann Arbor
Ch.1 Residual Neural Network Fundamentals - Overcoming Neural Network Degradation Through Residual Learning @ University of Michigan - Ann Arbor
Identity versus Projection Shortcuts - Overcoming Neural Network Degradation Through Residual Learning @ University of Michigan - Ann Arbor
Related
Residual Mapping
Shortcut’s technique for identity mapping
Dimension Matching and Projection Shortcuts
Bottleneck Residual Blocks
Empirical Performance of Identity vs. Projection Shortcuts
Dimension Matching and Projection Shortcuts
Bottleneck Residual Blocks
ImageNet Classification Results and Benchmarks
Deep ResNet Scaling with Bottleneck Blocks
Dimension Matching and Projection Shortcuts
Plain versus Residual Baseline Architectures
Empirical Performance of Identity vs. Projection Shortcuts
Shortcut’s technique for identity mapping
Dimension Matching and Projection Shortcuts
Inductive Bias of Residual Connections
ResNet Function Decomposition
Residual Connections Enable Deeper ResNet Training
As plain networks grow deeper, their training accuracy can paradoxically worsen, a phenomenon known as the ___ problem.
Order the steps carried out during a forward pass in a residual block to compute the target function f(x) from an input x.
Explain why reformulating the layers to learn the residual mapping g(x) makes approximating the identity function f(x) = x easier for the network compared to learning f(x) directly.
Match each concept from residual learning to its corresponding description.
Smaller ResNet Layer Responses and the Identity-Mapping Motivation
Dimension Matching and Projection Shortcuts
Learn After
Match each shortcut strategy for handling dimension changes to its corresponding implementation rule.
Order the operations performed to compute the output of a residual block with a projection shortcut when changing dimensions.
Explain why Option B achieves higher accuracy than Option A, and state which type of shortcut should be favored whenever dimensions match to minimize model size and computational complexity.
Match each mathematical or structural element of dimension matching to its specific role in residual networks.
Order the shortcut strategies by the total number of parameters they add to the network's shortcut connections, from fewest (least parameters) to most (greatest parameters).
Based on the evaluated shortcut strategies, evaluate whether switching from Option B to Option C is justified in terms of accuracy and model complexity, and explain why identity shortcuts are generally favored.
Match each shortcut evaluation strategy to its corresponding performance trade-off and empirical finding in residual networks.
When downsampling feature map sizes across stages in a residual network, the operation is performed with a stride of ___ .
Explain why the shortcut input cannot be combined directly with the residual output during this stage transition, and describe how applying a linear projection resolves this issue.
Empirical Performance of Identity vs. Projection Shortcuts