You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决train_test_split样本数不一致错误(ISED数据集场景)

Fixing Inconsistent Sample Count Error in train_test_split for Indian Spontaneous Expression Dataset

Let's break down why you're getting that Found input variables with inconsistent numbers of samples error and fix it step by step.

Root Cause of the Error

Your code is only saving the last image you read into the X variable, instead of collecting all 428 images. That's why when you load x later, its shape is (1080, 1920, 3) (the dimensions of a single image) instead of (428, 1080, 1920, 3) (428 images total). Meanwhile, your labels y correctly have 428 samples, leading to the mismatch.

Step-by-Step Fix

Here's the corrected version of your code, with key changes highlighted:

import cv2 as cv
import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split

# Load metadata
data = pd.read_excel('/content/drive/My Drive/ISED/details1.xlsx')

# Initialize a list to collect ALL images (this was missing before!)
images = []
count = 0

for path in data['img_path']:
    count += 1
    temp1 = path.replace("'", "")
    imgpath = "/content/drive/My Drive/ISED/" + temp1
    imgFile = cv.imread(imgpath)
    
    # Make sure the image was loaded correctly
    if imgFile is not None:
        images.append(imgFile)
        print(f"Loaded image {count}, shape: {imgFile.shape}")
    else:
        print(f"Warning: Could not load image {count} at path {imgpath}")

# Convert the list of images to a numpy array (shape: (428, 1080, 1920, 3))
X = np.array(images)

# Process labels
y = pd.get_dummies(data['emotion']).values  # Use .values instead of deprecated .as_matrix()

# Save the full dataset
np.save('fdataXISED', X)
np.save('flabelsISED', y)

print("\nPreprocessing Done")
print(f"Number of examples in dataset: {len(X)}")
print(f"Shape of X: {X.shape}")
print(f"Shape of y: {y.shape}")
print("X,y stored in fdataXISED.npy and flabelsISED.npy respectively")

# Load the saved data
x = np.load('./fdataXISED.npy')
y = np.load('./flabelsISED.npy')

print(f"\nLoaded data dtype: {x.dtype}")

# Preprocess the images (normalize)
x = x.astype(float)
# For 4D arrays, compute mean/std across samples, height, width axes (per color channel)
mean = np.mean(x, axis=(0,1,2))
std = np.std(x, axis=(0,1,2))
x -= mean
x /= std

print(f"After normalization - X shape: {x.shape}, y shape: {y.shape}")

# Now split the data correctly (samples match!)
X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.1, random_state=42)
X_train, X_valid, y_train, y_valid = train_test_split(X_train, y_train, test_size=0.1, random_state=42)

print(f"\nDataset split complete:")
print(f"Training set: {X_train.shape}, {y_train.shape}")
print(f"Validation set: {X_valid.shape}, {y_valid.shape}")
print(f"Test set: {X_test.shape}, {y_test.shape}")

Key Changes Made

  • Added an images list: This collects all 428 images instead of overwriting X with each single image.
  • Replaced .as_matrix() with .values: .as_matrix() is deprecated in pandas, so .values is the recommended way to get a numpy array from a DataFrame.
  • Fixed normalization logic: For 4D image arrays, we calculate mean and standard deviation across the sample, height, and width axes (keeping the channel axis separate) to properly normalize each color channel.
  • Added error checking: Checks if an image failed to load (in case of missing files or wrong paths), so you can debug any loading issues.

Why This Works

Now your x array will have the shape (428, 1080, 1920, 3) (428 samples, each 1080x1920 with 3 color channels), matching the 428 samples in y. This makes train_test_split work without sample count mismatches.

内容的提问来源于stack exchange,提问作者Nagesh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:37:48