You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Keras图像分类教程做多分类时遇样本数不匹配错误

Fixing Input-Target Sample Mismatch in Keras Multi-Class Image Classification

Hey there, let's tackle that sample count mismatch error you're hitting (ValueError: Input arrays should have the same number of samples as target arrays. Found 1600 input samples and 1800 target samples). The root issue comes from how you're generating your labels—you're using the hardcoded approach from the binary classification tutorial, which doesn't account for your 4-class setup and potentially unequal sample counts per class.

Here's how to fix it:

1. Generate labels based on actual sample counts (no hardcoding!)

Instead of guessing or hardcoding the number of samples per class, dynamically count the images in each class folder and generate labels to match. This ensures your label array length exactly matches the number of input samples you load.

import os
import numpy as np

# Define your dataset paths and class names
train_dir = "data/train"
class_names = ["class1", "class2", "class3", "class4"]

train_labels = []
total_train_samples = 0

# Iterate over each class to count samples and build labels
for class_index, class_name in enumerate(class_names):
    # Path to the current class's training images
    class_image_dir = os.path.join(train_dir, class_name)
    # Count how many images are in this class folder
    num_class_samples = len(os.listdir(class_image_dir))
    total_train_samples += num_class_samples
    # Add the class index as a label for every image in this class
    train_labels.extend([class_index] * num_class_samples)

# Convert the list to a numpy array for Keras
train_labels = np.array(train_labels)

2. Verify sample counts match

Before training, double-check that your input features and labels have the same number of samples:

# Assuming your bottleneck features are stored in bottleneck_features_train
print(f"Input samples: {bottleneck_features_train.shape[0]}")
print(f"Target samples: {len(train_labels)}")

These numbers should now be identical, eliminating the mismatch error.

3. Adjust for multi-class classification (critical extra step!)

Since you're moving from binary to multi-class classification, don't forget to update your model's loss function and output layer activation:

  • If you're using integer labels (like the train_labels array we generated above), use sparse_categorical_crossentropy as your loss function, and set the output layer activation to softmax:
    model.compile(
        optimizer="rmsprop",
        loss="sparse_categorical_crossentropy",
        metrics=["accuracy"]
    )
    
  • If you prefer one-hot encoded labels, convert your integer labels first and use categorical_crossentropy:
    from keras.utils import to_categorical
    
    train_labels_onehot = to_categorical(train_labels)
    model.compile(
        optimizer="rmsprop",
        loss="categorical_crossentropy",
        metrics=["accuracy"]
    )
    

This approach ensures your labels align perfectly with your input data, and your model is configured correctly for multi-class tasks.

内容的提问来源于stack exchange,提问作者Rocking chief

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 11:07:43