Keras中ResNet特征提取时Numpy图像数组维度错误的解决方法
Hey there, let's get that dimension mismatch sorted out—this is a super common gotcha when working with pre-trained CNNs in Keras!
First, let's break down the error: your ResNet50 model expects a 4-dimensional input array with the shape (number_of_samples, image_height, image_width, number_of_channels) (since it's built for RGB images, channels=3). But your current myarray is ending up as (324398, 1)—this usually happens because your loaded images don't have consistent dimensions, so NumPy can't stack them into a proper 4D tensor and instead creates a 2D array of objects.
Here's how to fix it step by step:
1. Standardize Image Size & Format
When loading your images, resize every one to match the input shape your ResNet50 model expects (default is 224x224 for ResNet50). Also, make sure all images are converted to 3-channel RGB (in case any are grayscale).
Update your image loading code like this:
from PIL import Image import numpy as np import glob from keras.applications.resnet50 import ResNet50, preprocess_input # Define the input shape expected by ResNet50 image_h, image_w = 224, 224 def load_image(fname): # Open the image and resize to standard dimensions img = Image.open(fname).resize((image_w, image_h)) # Convert grayscale images to RGB if needed if img.mode != 'RGB': img = img.convert('RGB') # Convert to numpy array return np.array(img) # Load and process all images filelist = glob.glob('dataset/*.jpg') myarray = np.array([load_image(fname) for fname in filelist])
2. Verify the Array Shape
After running the above code, check the shape of myarray with print(myarray.shape)—it should now be (324398, 224, 224, 3) (or whatever your sample count is). That's the 4D shape the model needs!
3. Add Proper Preprocessing (Critical!)
Since you're using ResNet50 with ImageNet weights, you need to apply the same preprocessing that the model was trained on. Update your feature extraction function to include this:
def extract_resnet(X): # Initialize the ResNet50 model without the top classification layer resnet_model = ResNet50(input_shape=(image_h, image_w, 3), weights='imagenet', include_top=False) # Apply ImageNet preprocessing to the input array X_processed = preprocess_input(X) # Extract features features_array = resnet_model.predict(X_processed) return features_array # Now run the extraction print(extract_resnet(myarray))
Why This Works
- By resizing all images to the same dimensions, NumPy can stack them into a proper 4D tensor instead of a messy 2D object array.
- Converting to RGB ensures we have the 3 channels ResNet50 expects.
- The
preprocess_inputfunction scales pixel values to match the range the model was trained on (this prevents weird predictions/errors from mismatched data distributions).
内容的提问来源于stack exchange,提问作者Juanvulcano

