如何解决train_test_split样本数不一致错误(ISED数据集场景)
Let's break down why you're getting that Found input variables with inconsistent numbers of samples error and fix it step by step.
Root Cause of the Error
Your code is only saving the last image you read into the X variable, instead of collecting all 428 images. That's why when you load x later, its shape is (1080, 1920, 3) (the dimensions of a single image) instead of (428, 1080, 1920, 3) (428 images total). Meanwhile, your labels y correctly have 428 samples, leading to the mismatch.
Step-by-Step Fix
Here's the corrected version of your code, with key changes highlighted:
import cv2 as cv import pandas as pd import numpy as np from sklearn.model_selection import train_test_split # Load metadata data = pd.read_excel('/content/drive/My Drive/ISED/details1.xlsx') # Initialize a list to collect ALL images (this was missing before!) images = [] count = 0 for path in data['img_path']: count += 1 temp1 = path.replace("'", "") imgpath = "/content/drive/My Drive/ISED/" + temp1 imgFile = cv.imread(imgpath) # Make sure the image was loaded correctly if imgFile is not None: images.append(imgFile) print(f"Loaded image {count}, shape: {imgFile.shape}") else: print(f"Warning: Could not load image {count} at path {imgpath}") # Convert the list of images to a numpy array (shape: (428, 1080, 1920, 3)) X = np.array(images) # Process labels y = pd.get_dummies(data['emotion']).values # Use .values instead of deprecated .as_matrix() # Save the full dataset np.save('fdataXISED', X) np.save('flabelsISED', y) print("\nPreprocessing Done") print(f"Number of examples in dataset: {len(X)}") print(f"Shape of X: {X.shape}") print(f"Shape of y: {y.shape}") print("X,y stored in fdataXISED.npy and flabelsISED.npy respectively") # Load the saved data x = np.load('./fdataXISED.npy') y = np.load('./flabelsISED.npy') print(f"\nLoaded data dtype: {x.dtype}") # Preprocess the images (normalize) x = x.astype(float) # For 4D arrays, compute mean/std across samples, height, width axes (per color channel) mean = np.mean(x, axis=(0,1,2)) std = np.std(x, axis=(0,1,2)) x -= mean x /= std print(f"After normalization - X shape: {x.shape}, y shape: {y.shape}") # Now split the data correctly (samples match!) X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.1, random_state=42) X_train, X_valid, y_train, y_valid = train_test_split(X_train, y_train, test_size=0.1, random_state=42) print(f"\nDataset split complete:") print(f"Training set: {X_train.shape}, {y_train.shape}") print(f"Validation set: {X_valid.shape}, {y_valid.shape}") print(f"Test set: {X_test.shape}, {y_test.shape}")
Key Changes Made
- Added an
imageslist: This collects all 428 images instead of overwritingXwith each single image. - Replaced
.as_matrix()with.values:.as_matrix()is deprecated in pandas, so.valuesis the recommended way to get a numpy array from a DataFrame. - Fixed normalization logic: For 4D image arrays, we calculate mean and standard deviation across the sample, height, and width axes (keeping the channel axis separate) to properly normalize each color channel.
- Added error checking: Checks if an image failed to load (in case of missing files or wrong paths), so you can debug any loading issues.
Why This Works
Now your x array will have the shape (428, 1080, 1920, 3) (428 samples, each 1080x1920 with 3 color channels), matching the 428 samples in y. This makes train_test_split work without sample count mismatches.
内容的提问来源于stack exchange,提问作者Nagesh

