An object detector proposes regions of wildly different sizes — a tall pedestrian box, a wide car box — and then has to feed each one to a dense layer that demands a fixed input size. ROI pooling solves it by carving any region into a fixed grid of bins and taking the maximum in each.
Task: write roi_pool(feature_map, roi, out_size) returning an out_size × out_size grid.
roi is [row_start, col_start, row_end, col_end], with both ends inclusive. So its height is row_end - row_start + 1 and its width col_end - col_start + 1.
For output cell (i, j), the bin covers rows:
and the columns follow the same formula with j and w. The lower bound is inclusive, the upper exclusive. Take the maximum feature value over that block.
floor on the start and ceil on the end is the standard trick for dividing a region that doesn't divide evenly: every bin gets at least one row and column, and neighbouring bins may share a row where the split falls mid-pixel.out_size is never larger than the ROI's height or width, so no bin is ever empty.The point of all this is that the output size is fixed by out_size alone, no matter the region's dimensions. That's what lets one trained dense head score thousands of differently-shaped proposals from a single pass over the feature map — the architectural idea that made Fast R-CNN fast.