基于Keras图像分类教程做多分类时遇样本数不匹配错误
Hey there, let's tackle that sample count mismatch error you're hitting (ValueError: Input arrays should have the same number of samples as target arrays. Found 1600 input samples and 1800 target samples). The root issue comes from how you're generating your labels—you're using the hardcoded approach from the binary classification tutorial, which doesn't account for your 4-class setup and potentially unequal sample counts per class.
Here's how to fix it:
1. Generate labels based on actual sample counts (no hardcoding!)
Instead of guessing or hardcoding the number of samples per class, dynamically count the images in each class folder and generate labels to match. This ensures your label array length exactly matches the number of input samples you load.
import os import numpy as np # Define your dataset paths and class names train_dir = "data/train" class_names = ["class1", "class2", "class3", "class4"] train_labels = [] total_train_samples = 0 # Iterate over each class to count samples and build labels for class_index, class_name in enumerate(class_names): # Path to the current class's training images class_image_dir = os.path.join(train_dir, class_name) # Count how many images are in this class folder num_class_samples = len(os.listdir(class_image_dir)) total_train_samples += num_class_samples # Add the class index as a label for every image in this class train_labels.extend([class_index] * num_class_samples) # Convert the list to a numpy array for Keras train_labels = np.array(train_labels)
2. Verify sample counts match
Before training, double-check that your input features and labels have the same number of samples:
# Assuming your bottleneck features are stored in bottleneck_features_train print(f"Input samples: {bottleneck_features_train.shape[0]}") print(f"Target samples: {len(train_labels)}")
These numbers should now be identical, eliminating the mismatch error.
3. Adjust for multi-class classification (critical extra step!)
Since you're moving from binary to multi-class classification, don't forget to update your model's loss function and output layer activation:
- If you're using integer labels (like the
train_labelsarray we generated above), usesparse_categorical_crossentropyas your loss function, and set the output layer activation tosoftmax:model.compile( optimizer="rmsprop", loss="sparse_categorical_crossentropy", metrics=["accuracy"] ) - If you prefer one-hot encoded labels, convert your integer labels first and use
categorical_crossentropy:from keras.utils import to_categorical train_labels_onehot = to_categorical(train_labels) model.compile( optimizer="rmsprop", loss="categorical_crossentropy", metrics=["accuracy"] )
This approach ensures your labels align perfectly with your input data, and your model is configured correctly for multi-class tasks.
内容的提问来源于stack exchange,提问作者Rocking chief

