ResNet Adaptation in Faster R-CNN
Applying ResNet-50 or ResNet-101 within Faster R-CNN requires replacing the hidden fully connected layers used by architectures such as VGG-16. Under the Networks on Conv feature maps (NoC) approach, conv1 through conv4_x compute shared full-image feature maps whose effective cumulative stride is at most 16 pixels; in ResNet-101, these stages contain 91 convolutional layers. Both the Region Proposal Network, which generates 300 proposals, and the Fast R-CNN detector use these shared features. RoI pooling is inserted before conv5_1, and conv5_x is then applied separately to each RoI before the sibling classification and bounding-box regression branches. To limit memory use during fine-tuning, Batch Normalization means and variances precomputed on the ImageNet training set are held fixed. The Batch Normalization layers therefore operate as fixed affine transformations rather than updating their statistics from detection-training mini-batches.
0
1
Tags
Prep Sessions
Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Ch.4 Residual Network Experiments and Applications - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Object Detection and Localization using Residual Networks - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Learn After
Match each component or layer group of the Faster R-CNN ResNet adaptation to its specific architectural role.
Place the processing stages of an image through a ResNet-101 Faster R-CNN architecture in the correct sequential order.
Explain how the Faster R-CNN ResNet adaptation resolves this memory consumption issue, detailing the specific treatment of Batch Normalization layers and their resulting functional behavior during fine-tuning.