Keras实现Segnet报错‘新数组总大小必须不变’的解决方法咨询
Let's break down why this error pops up and how to fix it without forcing your images into a 256x256 box:
Why the Error Happens
SegNet’s core is a symmetric encoder-decoder structure that relies on max pooling (downsampling) and unpooling (upsampling). Each pooling step halves the image’s spatial dimensions, so your input size needs to be divisible by (2^n) — where (n) is the number of pooling layers (usually 4 or 5 in standard SegNet). When your original images don’t meet this requirement, the unpooling step tries to upsample to a dimension that doesn’t align with the encoder’s feature maps, triggering the array size mismatch error.
Solutions to Preserve Original Image Dimensions
1. Pad Images to Meet Dimension Requirements (Recommended)
Instead of resizing, pad your images to the nearest multiple of 16 or 32 (depending on your pooling layer count) while keeping their original aspect ratio. After inference, you can crop the output back to the original size to avoid distortion.
Here’s a practical example using OpenCV:
import cv2 import numpy as np def pad_to_divisible(img, divisor=16): h, w = img.shape[:2] # Calculate padding needed to make dimensions divisible by the target number pad_h = (divisor - h % divisor) % divisor pad_w = (divisor - w % divisor) % divisor # Split padding evenly between top/bottom and left/right to keep center aligned pad_top = pad_h // 2 pad_bottom = pad_h - pad_top pad_left = pad_w // 2 pad_right = pad_w - pad_left # Use mirror padding to avoid edge artifacts (better than zero padding for segmentation) padded_img = cv2.copyMakeBorder(img, pad_top, pad_bottom, pad_left, pad_right, cv2.BORDER_REFLECT) return padded_img, (pad_top, pad_bottom, pad_left, pad_right) # Usage during preprocessing original_img = cv2.imread("your_image.jpg") padded_img, padding_params = pad_to_divisible(original_img, divisor=16) # After inference, crop the output mask back to original size output_mask = model(padded_img_tensor) cropped_mask = output_mask[padding_params[0]:-padding_params[1], padding_params[2]:-padding_params[3]]
2. Modify SegNet’s Downsampling/Upsampling Layers
If you don’t want to adjust your raw data, tweak the model to use strided convolutions instead of max pooling for downsampling, and transposed convolutions for upsampling. This gives you more control over output dimensions.
For example, replace a max pooling layer:
# Original max pooling (returns indices for unpooling) nn.MaxPool2d(2, return_indices=True) # Replace with strided convolution for downsampling nn.Conv2d(in_channels, out_channels, kernel_size=3, stride=2, padding=1)
And in the decoder, use a transposed convolution to match the dimension:
# Replace unpooling with transposed convolution nn.ConvTranspose2d(in_channels, out_channels, kernel_size=3, stride=2, padding=1, output_padding=1)
Note: This changes the original SegNet architecture, so you’ll need to validate if segmentation performance stays consistent.
3. Debug Layer Dimensions Step-by-Step
Sometimes the error stems from a mismatch between your input, model layers, or label sizes. Print the shape of each layer’s output to pinpoint where the dimension breakdown happens:
from torchsummary import summary # Assuming your model is stored in 'segnet_model' summary(segnet_model, input_size=(3, original_h, original_w))
Or manually trace intermediate outputs:
dummy_input = torch.randn(1, 3, original_h, original_w) x = dummy_input for idx, layer in enumerate(segnet_model.encoder): x = layer(x) print(f"Encoder layer {idx} output shape: {x.shape}") # Repeat this for decoder layers to spot mismatches
4. Use Adaptive Pooling Layers
Swap fixed max pooling layers with adaptive pooling to let the model handle arbitrary input sizes:
# Replace fixed max pooling with adaptive pooling (target size matches decoder input) nn.AdaptiveMaxPool2d((target_h, target_w))
You’ll need to adjust the decoder’s upsampling layers to match the target dimensions, but this gives you full flexibility with input sizes.
内容的提问来源于stack exchange,提问作者Haroon S.

