CNN层输出体积计算:公式((n + 2p - f / s) + 1)使用是否正确?
Short answer: No — the formula you've written has a critical order-of-operations error that makes it invalid for calculating CNN layer output dimensions.
The Correct Formula
The standard formula for computing the output height/width of a convolutional or pooling layer is:
output_size = floor((n + 2p - f) / s) + 1
Where:
n: Input height or width dimensionp: Padding added to each side of the inputf: Size of the convolutional/pooling kernelfloor(): Ensures we get an integer (since you can't have a fraction of a pixel in output)
What's Wrong With Your Formula?
Your formula ((n + 2p - f / s) + 1) incorrectly applies division only to the filter size f, instead of dividing the entire (n + 2p - f) term by the stride s. This breaks the calculation entirely, as we'll see with examples from your provided layers.
Let's Verify With Your Layer Data
Let's use the layers you listed to demonstrate why the correct formula works and yours doesn't:
1. POOL 1 Layer
- Input size (from CONV 1 output):
1242 × 378 - Output size:
621 × 189 - Assume standard max pooling parameters:
f=2,s=2,p=0- Correct formula:
(1242 + 0 - 2) / 2 + 1 = 1240/2 +1 = 620 +1 =621(matches the given output perfectly) - Your formula:
((1242 +0 - 2/2)+1) = (1242 -1)+1=1242(way off, gives the input size instead of the pooled output)
- Correct formula:
2. CONV 2 Layer
- Input size (from POOL 1 output):
621 ×189 - Output size:
310 ×94 - Reverse-engineering parameters: Let's say
s=2,p=0,f=3- Correct formula:
(621 +0 -3)/2 +1 =618/2 +1=309+1=310(matches the output) - Your formula:
((621+0 -3/2)+1)=621-1.5+1=620.5(non-integer result, which is impossible for a layer output)
- Correct formula:
3. POOL 2 Layer
- Input size (from CONV 2 output):
310 ×94 - Output size:
78 ×24 - Reverse-engineering parameters:
f=4,s=4,p=1- Correct formula:
(310 +2*1 -4)/4 +1=(312-4)/4 +1=308/4+1=77+1=78(matches) - Your formula:
((310+2 -4/4)+1)=312-1+1=312(completely incorrect, gives input size again)
- Correct formula:
Key Takeaway
The division in the formula must apply to the entire (n + 2p - f) term, not just the filter size. Using the correct order of operations ensures you get valid, integer output dimensions that match real CNN layer behavior.
内容的提问来源于stack exchange,提问作者user9516512

