如何计算ROI Align处理后的特征图?(基于Mask RCNN论文)
Hey there! Let's walk through exactly how to compute ROI Align outputs using bilinear interpolation, step by step. I'll use a concrete example to make this tangible, since you mentioned struggling with the slide scenario.
First, Let's Recap the Core Steps of ROI Align
ROI Align fixes the misalignment issue of ROI Pooling by preserving floating-point coordinates and using bilinear interpolation to sample values. Here's the high-level workflow:
- Map the ROI to the feature space: Take your original image ROI and scale it down to match the feature map's resolution (e.g., if your backbone uses 16x downsampling, divide the ROI's x/y coordinates by 16).
- Divide the mapped ROI into fixed-size bins: If you want a 2x2 output feature map, split the scaled ROI into 2x2 equal-sized bins.
- Sample points within each bin: For each bin, pick 4 regularly spaced sampling points (this is the standard approach from the Mask R-CNN paper).
- Bilinear interpolation for each sample: Calculate the value of each floating-point sample point using its surrounding 4 integer-coordinate pixels in the feature map.
- Average the samples per bin: Take the mean of the 4 interpolated values in each bin to get the final value for that position in the output feature map.
Let's Break Down the Bilinear Interpolation Math
For any floating-point coordinate (x, y) in the feature map:
- Find the 4 nearest integer coordinates:
(x1, y1) = (floor(x), floor(y)),(x2, y2) = (x1+1, y1+1) - The interpolated value is calculated as:
Wherevalue = (x2 - x) * (y2 - y) * f(x1,y1) + (x - x1) * (y2 - y) * f(x2,y1) + (x2 - x) * (y - y1) * f(x1,y2) + (x - x1) * (y - y1) * f(x2,y2)f(a,b)is the pixel value of the feature map at integer coordinate(a,b).
Concrete Example to Match Your Slide Scenario
Let’s say your scaled ROI (on the feature map) is a 4x4 grid with these values:
[[1, 2, 3, 4], [5, 6, 7, 8], [9,10,11,12], [13,14,15,16]]
We want to output a 2x2 feature map, so split this into 4 2x2 bins. Let's compute the top-left bin first:
Step 1: Define the bin and its sample points
The top-left bin covers feature map coordinates x ∈ [0, 2) and y ∈ [0, 2). We pick 4 sample points at relative positions (0.5,0.5), (0.5,1.5), (1.5,0.5), (1.5,1.5) within the bin.
Step 2: Calculate each sample with bilinear interpolation
- Sample (0.5, 0.5):
Surrounding pixels are(0,0)=1,(1,0)=2,(0,1)=5,(1,1)=6value = 0.5*0.5*1 + 0.5*0.5*2 + 0.5*0.5*5 + 0.5*0.5*6 = (1+2+5+6)/4 = 3.5 - Sample (0.5, 1.5):
Surrounding pixels are(0,1)=5,(1,1)=6,(0,2)=9,(1,2)=10value = 0.5*0.5*5 + 0.5*0.5*6 + 0.5*0.5*9 + 0.5*0.5*10 = (5+6+9+10)/4 = 7.5 - Sample (1.5, 0.5):
Surrounding pixels are(1,0)=2,(2,0)=3,(1,1)=6,(2,1)=7value = 0.5*0.5*2 + 0.5*0.5*3 + 0.5*0.5*6 + 0.5*0.5*7 = (2+3+6+7)/4 = 4.5 - Sample (1.5, 1.5):
Surrounding pixels are(1,1)=6,(2,1)=7,(1,2)=10,(2,2)=11value = 0.5*0.5*6 + 0.5*0.5*7 + 0.5*0.5*10 + 0.5*0.5*11 = (6+7+10+11)/4 = 8.5
Step 3: Average the samples for the bin
The top-left bin's output value is (3.5 + 7.5 + 4.5 + 8.5)/4 = 6.
Repeat this process for the remaining 3 bins, and you’ll get your final 2x2 ROI Align output.
Key Difference from ROI Pooling
Remember: ROI Pooling would round the floating-point bin coordinates to integers, leading to misalignment between the ROI and feature map. ROI Align skips this rounding and uses bilinear interpolation instead, which is why it produces much more accurate results for mask prediction in Mask R-CNN.
内容的提问来源于stack exchange,提问作者Abhijeet Sharma

