Per-Class Localization and RoI-Centric Training
For the ImageNet Localization (LOC) task, per-class regression (PCR) learns a separate bounding-box regressor for each of the 1,000 object categories. The category-agnostic proposal network is replaced by a per-class RPN ending in two sibling convolutional layers: a 1,000-dimensional binary logistic classification layer that predicts class presence and a -dimensional regression layer that predicts coordinate offsets relative to translation-invariant anchor boxes. Because ImageNet images typically contain one dominant object, their highly overlapping region proposals produce nearly identical RoI-pooled features, reducing sample variance and hindering stochastic image-centric training. The framework therefore uses an RoI-centric R-CNN pipeline. During training, the top 200 proposals from the per-class RPN for ground-truth classes are cropped, warped to pixels, and sampled in mini-batches of 256 RoIs. During inference, the per-class RPN extracts the top 200 proposals for each predicted class, after which the R-CNN scores and regresses them. A single ResNet-101 model achieves 10.6% top-5 localization error on the ImageNet validation set, while an ensemble achieves 9.0% on the test set and first place in the ILSVRC 2015 localization challenge.
0
1
Tags
Prep Sessions
Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Ch.4 Residual Network Experiments and Applications - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Object Detection and Localization using Residual Networks - Foundational Deep Learning Architectures: Transformers and Residual Networks @ University of Michigan - Ann Arbor
Learn After
Match each architectural component or parameter from the ImageNet per-class localization framework to its corresponding role.
Order the stages of the ResNet localization pipeline during inference from first to last.
Based on the behavior of region proposals in ImageNet, explain why image-centric training causes stochastic training to stall, and state the pipeline change adopted to overcome this issue.