样本数不匹配ValueError报错及图像与标签同步洗牌需求
sklearn.utils.shuffle Let's break down what's causing this error and fix it step by step.
What's Going Wrong?
The error ValueError: Found input variables with inconsistent numbers of samples: [3058, 3777] tells us your image matrix M has 3058 samples, but your label array has 3777. This mismatch happens because of a few key issues in your code:
Variable Name Typos & Incorrect Folder Selection
You defined your folder list asfolder, but tried to concatenatefolders[1]tofolders[8](note the extra 's'). This likely pulled in the wrong data. Worse, you skipped the first folder (ALB, index 0) which should contribute 1720 samples, but includedtest images(index 8) which shouldn't be part of your training labels—this threw off the total sample count completely.Manual Label Counting Error
You hardcoded the label slices to 3777, but your actual concatenated images didn't add up to that number because of the folder selection mistake. Manual counting is prone to errors—always derive labels from the actual data you load.Uninitialized List
L
You calledL.append(len(listing))but never initializedL = [], which would throw an error before even reaching the shuffle step (maybe you fixed this but didn't include it in your code snippet).
Step-by-Step Solution
Here's a revised version of your code that fixes these issues and ensures your data and labels are perfectly aligned:
import numpy as np import os import cv2 from sklearn.utils import shuffle # Define only the training folders (exclude 'test images' since it has no labels) folder_names = ['ALB', 'BET', 'DOL', 'LAG', 'NoF', 'OTHER', 'SHARK', 'YFT'] data_path = r'D:\ncfm\train' image_arrays = [] sample_counts = [] for class_name in folder_names: # Get full path to the class folder class_folder = os.path.join(data_path, class_name) # List all image files in the folder image_files = os.listdir(class_folder) # Load, convert to grayscale, flatten images # Use cv2.IMREAD_GRAYSCALE to ensure consistent grayscale conversion class_images = np.array([ np.array(cv2.imread(os.path.join(class_folder, img), cv2.IMREAD_GRAYSCALE)).flatten() for img in image_files ]) image_arrays.append(class_images) sample_counts.append(len(class_images)) # Concatenate all class image arrays into one matrix M = np.concatenate(image_arrays) # Generate labels dynamically based on actual sample counts total_samples = sum(sample_counts) label = np.zeros(total_samples, dtype=int) current_index = 0 # Assign class labels (1 to 8 as in your original code) for class_idx, count in enumerate(sample_counts, start=1): label[current_index:current_index + count] = class_idx current_index += count # Now shuffle the data and labels safely data, Label = shuffle(M, label, random_state=2) # Verify the alignment (optional but recommended) print(f"Data shape: {data.shape}, Label shape: {Label.shape}")
Key Improvements
- Dynamic Label Generation: Labels are created based on the actual number of images loaded, so no more manual counting errors.
- Correct Folder Selection: We only include the 8 labeled training folders, skipping
test imageswhich doesn't belong in your training set. - Consistent Grayscale Conversion: Using
cv2.IMREAD_GRAYSCALEensures every image is loaded as grayscale consistently, avoiding unexpected shape issues. - Clear Variable Names: Renamed variables like
sample_countsinstead ofLfor readability, anddata_pathinstead of reusingPath(which overwrites thepathlib.Pathimport).
How to Verify
After running the code, the print statement should output something like Data shape: (3777, X), Label shape: (3777,), confirming the sample counts match. If they don't, double-check that each folder has the number of images you expect (cross-reference with sample_counts).
内容的提问来源于stack exchange,提问作者SREE LAKSHMI

