阅读Chollet《Python深度学习》的疑问:连续Conv2D层中7×7窗口来源
Great question—let's break this down step by step, starting with where that 7x7 number comes from.
Where Does the 7x7 Come From?
First, let's assume your input is the standard 28x28 MNIST digit image (since that's the example François Chollet uses in Python Deep Learning for this kind of experiment). Looking at your layer outputs:
conv2d_4takes 28x28 input and outputs 26x26conv2d_5takes 26x26 and outputs 24x24conv2d_6takes 24x24 and outputs 22x22
Each of these Conv2D layers uses a 3x3 kernel, stride=1, and no padding (valid padding)—that's why the spatial size decreases by 2 each time (output size = input size - kernel size + 1 → 28-3+1=26, etc.).
Now, let's calculate the receptive field of a single pixel in the final conv2d_6 layer. The receptive field is the area of the original input that contributes to that pixel:
- A pixel in
conv2d_4depends on a 3x3 patch of the original input. - A pixel in
conv2d_5depends on a 3x3 patch ofconv2d_4's output. Each of thoseconv2d_4pixels covers 3x3 input, so the total input area here is 3 + (3-1) = 5x5 (we add 2 because each edge of the kernel expands the field by 1 pixel on each side). - A pixel in
conv2d_6depends on a 3x3 patch ofconv2d_5's output. Following the same logic, we add another 2 pixels to each dimension: 5 + 2 = 7x7.
Put simply: three consecutive 3x3 conv layers (with stride 1, no padding) result in each final feature pixel only seeing a 7x7 window of the original 28x28 digit.
Why This Hurts Digit Classification
Digits like 0, 8, or 9 rely on their full shape to be recognized—you can't tell a 0 from an 8 by looking at just a 7x7 patch. Chollet's point is that without pooling layers (like MaxPooling), your high-level features never get a chance to "see" the entire digit.
MaxPooling layers (usually 2x2 with stride 2) double the receptive field each time they're used. For example, if you added a MaxPooling layer after each conv layer, the receptive field would grow much faster, and by the final layer, your features would cover the entire 28x28 input—critical for accurate digit classification.
内容的提问来源于stack exchange,提问作者MeiNan Zhu

