为何Keras ResNet50首个池化输出为55x55,而论文中是56x56?
I've run into this exact issue before—using a pre-trained ResNet50 as a U-Net encoder and hitting a size mismatch between encoder features and decoder upsampled maps is super frustrating. Let's break down why this is happening and how to fix it.
Why the 55x55 vs 56x56 Mismatch?
The root cause is likely the padding setting in the first max pooling layer (pool1) of your ResNet50 implementation. In older versions of Keras Applications, the pool1 layer uses padding='valid' by default instead of padding='same' as specified in the original ResNet paper.
Here's the math behind the size difference:
- Input: 224x224x3
- After the initial 7x7 conv (stride=2, with proper padding) → 112x112x64
- With
padding='valid'in 3x3 max pool (stride=2):(112 - 3)/2 + 1 = 55→ 55x55x64 - With
padding='same':112//2 = 56→ 56x56x64 (matches the paper's expected output)
This 1-pixel gap breaks your skip connection when the decoder upsamples to 56x56.
Solution 1: Fix the ResNet50 Encoder's Pooling Layer
The cleanest fix is to modify the pool1 layer to use padding='same' while retaining the pre-trained weights. Here's how to do it with TensorFlow Keras:
import tensorflow as tf from tensorflow.keras.applications.resnet50 import ResNet50 from tensorflow.keras.layers import MaxPooling2D from tensorflow.keras.models import clone_model # Load the base pre-trained ResNet50 (without top classification layers) base_model = ResNet50(include_top=False, input_shape=(224, 224, 3), weights='imagenet') base_model.layers.pop() # Remove the final average pooling layer # Function to update the pool1 layer's padding setting def update_pool1_config(layer): if layer.name == 'pool1': # Replace with same-padding max pooling to match paper specs return MaxPooling2D((3, 3), strides=(2, 2), padding='same', name='pool1') return layer # Clone the model with the modified pool1 layer fixed_encoder = clone_model(base_model, clone_function=update_pool1_config) # Load original pre-trained weights (pool1 has no trainable weights, so this is safe) fixed_encoder.set_weights(base_model.get_weights()) # Verify the fix—check the summary for conv2_x outputs (now 56x56) fixed_encoder.summary()
Now your encoder's first residual block output will be 56x56, perfectly matching the decoder's upsampled features for concatenation.
Solution 2: Adjust Feature Map Sizes (Quick Fix)
If you don't want to rebuild the encoder, you can adjust either the encoder's output or the decoder's output to resolve the mismatch:
Option A: Upsample Encoder Features
Resize the 55x55 encoder feature map to 56x56 using bilinear interpolation:
from tensorflow.keras.layers import Lambda import tensorflow as tf # Get the first residual block output (e.g., from the base model) encoder_feature = base_model.get_layer('conv2_block3_out').output # Resize to match decoder's 56x56 size resized_feature = Lambda(lambda x: tf.image.resize(x, (56, 56), method='bilinear'))(encoder_feature)
Option B: Crop Decoder Features
Crop the 56x56 decoder feature map to 55x55:
from tensorflow.keras.layers import Cropping2D # Get your decoder's upsampled 56x56 feature map decoder_feature = ... # Replace with your decoder upsampling output # Crop 1 pixel from the bottom and right edges cropped_decoder = Cropping2D(cropping=((0, 1), (0, 1)))(decoder_feature)
Final Notes
Solution 1 is the most aligned with the original ResNet paper and will prevent future size mismatches in deeper skip connections (like 28x28, 14x14). The quick fixes work for temporary solutions, but modifying the encoder is the long-term clean approach.
内容的提问来源于stack exchange,提问作者Wei Zheng

