MLP模型训练遇ValueError:数据基数不匹配问题求助
MLP训练时数据样本数不匹配问题
我在训练MLP模型时遇到数据样本数不匹配的报错,怀疑是数据拆分和SMOTE过采样操作导致,但找不到具体原因。
代码及输出
数据拆分与SMOTE代码
# 将数据拆分为训练集、验证集和测试集 X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=seed) X_dev, X_test, y_dev, y_test = train_test_split(X_test, y_test, test_size=0.5, random_state=seed) # 打印训练集、验证集和测试集的形状 print("Training set shape:", X_train.shape, y_train.shape) print("Validation set shape:", X_dev.shape, y_dev.shape) print("Test set shape:", X_test.shape, y_test.shape) # SMOTE过采样 from imblearn.over_sampling import SMOTE smote = SMOTE(random_state=1) X_train_resampled, y_train_resampled = smote.fit_resample(X_train, y_train)
数据集形状输出
Training set shape:(7000, 5) (7000,) Validation set shape: (1500, 5) (1500,) Test set shape: (1500, 5) (1500,) Resampled training set shape: (13536, 5) (13536,) Training set shape: (7000, 5) (7000,) Validation set shape: (1530, 5) (1530,) Testing set shape: (1500, 5) (1500,)
MLP模型代码
from sklearn.neural_network import MLPClassifier from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score from sklearn import neural_network model = neural_network.MLPClassifier(random_state=1) from keras.models import Sequential from keras.layers import Dense,Dropout # 设置随机种子 import random random.seed(1) import numpy as np # numpy依赖random np.random.seed(1) import tensorflow as tf # tensorflow依赖numpy tf.random.set_seed(1) # 第一层 model.add(Dense(4,input_dim=4,activation = "relu")) # 注释:原注释写的8 neurons for 2nd layer,这里可能笔误 model.add(Dropout(0.1)) # 注释:Out of 8 neurons Drop 1 neuron randomly per epoch. 正则化提升泛化能力 # 输出层 model.add(Dense(1,activation = "sigmoid")) # 二分类任务,输出1个节点用sigmoid,也可以用Dense(2,activation="softmax") # model.add(Dense(n,activation="softmax")) # 多分类任务用这个 model.compile(loss="binary_crossentropy",optimizer="adam",metrics=["accuracy"]) # loss是训练过程中的损失,metrics是 epoch结束时评估指标 # 注释:正负样本不对称时不要用accuracy作为唯一指标 # 注释:adam是自适应学习率优化器,学习率依赖动量计算 b = 5 h = model.fit(X_train,Y_train,batch_size=b,epochs=100)
错误信息
ValueError Traceback (most recent call last) <ipython-input-328-208a2cc838ae> in <module> 1 b = 5 ----> 2 h = model.fit(X_train,Y_train,batch_size=b,epochs=100) 1 frames /usr/local/lib/python3.9/dist-packages/keras/engine/data_adapter.py in _check_data_cardinality(data) 1846 ) 1847 msg += "Make sure all arrays contain the same number of samples." -> 1848 raise ValueError(msg) 1849 1850 ValueError: Data cardinality is ambiguous: x sizes: 7000 y sizes: 7500 Make sure all arrays contain the same number of samples.
我尝试调整过数据拆分方式,但始终无法解决该问题,希望能得到帮助和建议,谢谢!
问题原因及解决方案
核心问题
- 变量名大小写错误:训练代码中用了
Y_train,但之前数据拆分得到的是小写的y_train,如果Y_train是其他地方定义的变量,必然导致样本数不匹配。 - SMOTE后未使用重采样数据:已经生成了过采样后的
X_train_resampled和y_train_resampled,但训练时仍用原始X_train,既浪费过采样操作,也可能因其他变量污染导致Y样本数异常。 - 模型定义冲突:先实例化sklearn的
MLPClassifier,又用Keras的Sequential往该模型加层,导致对象混乱,Keras模型需单独初始化。 - 输入维度不匹配:X数据集是5维(如
X_train.shape=(7000,5)),但第一层设置input_dim=4,后续也会引发错误。
修复步骤
- 统一变量名:确保训练时X和Y变量对应,比如用
y_train或重采样后的y_train_resampled,严格注意大小写一致。 - 使用重采样训练数据:若用SMOTE解决类别不平衡,训练时传入
X_train_resampled和y_train_resampled。 - 修正模型初始化:Keras模型单独初始化,不要复用sklearn的模型变量:
model = Sequential() model.add(Dense(4, input_dim=5, activation="relu")) # input_dim与X的特征数保持一致,改为5 model.add(Dropout(0.1)) model.add(Dense(1, activation="sigmoid"))
- 检查变量污染:排查代码中是否有其他地方修改过
y_train或定义了额外的Y_train变量,导致样本数不一致。
修正后的训练代码示例
# 初始化Keras模型 model = Sequential() model.add(Dense(4, input_dim=5, activation="relu")) model.add(Dropout(0.1)) model.add(Dense(1, activation="sigmoid")) model.compile(loss="binary_crossentropy",optimizer="adam",metrics=["accuracy"]) # 使用重采样后的训练数据 b = 5 h = model.fit(X_train_resampled, y_train_resampled, batch_size=b, epochs=100)
额外注意点
- 验证集和测试集禁止过采样,仅对训练集做过采样,避免数据泄露。
内容的提问来源于stack exchange,提问作者wayne Wen
相关产品推荐
相关产品推荐

