解决train_test_split报错:输入变量样本数不一致[40000,10000]
问题描述
执行代码行X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0)时出现错误:
Error: ValueError: Found input variables with inconsistent numbers of samples: [40000, 10000]
推测是向量化操作后X的数组尺寸变化,导致与y的样本数不匹配。以下是运行输出和代码:
运行输出
(10000, 4) (10000,) (40000, 1500) (10000,)
完整代码
import pandas as pd from sklearn.model_selection import train_test_split from sklearn.feature_extraction.text import CountVectorizer # 导入数据集: dataset = pd.read_excel(r"C:\Users\HPS1RT\Downloads\test\Safety_Prediction.xlsx", nrows=10000) dataset[["Safety"]] *= 1 # 为X和y变量赋值: X = dataset.iloc[:, :-1].values y = dataset.iloc[:, 4].values print(X.shape) print(y.shape) # 向量化 vectorizer = CountVectorizer(max_features=1500, min_df=5, max_df=0.7) X = vectorizer.fit_transform(X.ravel()).toarray() from sklearn.feature_extraction.text import TfidfTransformer tfidfconverter = TfidfTransformer() X = tfidfconverter.fit_transform(X).toarray() print(X.shape) print(y.shape) # 将数据集拆分为随机训练集和测试集: X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0) print(X_train) print(y_train)
问题原因与解决方法
核心问题
你用X.ravel()把原本(10000,4)的二维数组拉平成了40000个元素的一维数组,CountVectorizer会把每个元素当成独立样本处理,导致X变成(40000,1500),和y的10000个样本数量完全不匹配,从而触发错误。
解决思路
根据你的需求选择对应方案:
方案1:合并多列文本为单个样本特征(最常见需求)
如果每行的4个列属于同一样本的不同文本属性,需要先将它们合并成一个字符串,再做向量化:
# 替换原代码中X赋值和向量化的部分 import pandas as pd from sklearn.model_selection import train_test_split from sklearn.feature_extraction.text import CountVectorizer from sklearn.feature_extraction.text import TfidfTransformer dataset = pd.read_excel(r"C:\Users\HPS1RT\Downloads\test\Safety_Prediction.xlsx", nrows=10000) dataset[["Safety"]] *= 1 # 将4列文本合并为单个字符串,用空格分隔 X = dataset.iloc[:, :-1].apply(lambda row: ' '.join(row.astype(str)), axis=1).values y = dataset.iloc[:, 4].values print(X.shape) # 输出(10000,) print(y.shape) # 输出(10000,) # 向量化(无需使用ravel) vectorizer = CountVectorizer(max_features=1500, min_df=5, max_df=0.7) X = vectorizer.fit_transform(X).toarray() tfidfconverter = TfidfTransformer() X = tfidfconverter.fit_transform(X).toarray() print(X.shape) # 输出(10000, 1500),与y样本数一致 # 拆分数据集 X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0)
方案2:保留多列独立特征后拼接
如果需要让4个字段分别生成独立的文本特征再合并:
# 替换原代码中向量化的部分 import numpy as np import pandas as pd from sklearn.model_selection import train_test_split from sklearn.feature_extraction.text import CountVectorizer from sklearn.feature_extraction.text import TfidfTransformer dataset = pd.read_excel(r"C:\Users\HPS1RT\Downloads\test\Safety_Prediction.xlsx", nrows=10000) dataset[["Safety"]] *= 1 X = dataset.iloc[:, :-1].values y = dataset.iloc[:, 4].values print(X.shape) print(y.shape) # 对每一列单独做向量化 vectorizer = CountVectorizer(max_features=1500, min_df=5, max_df=0.7) X_vecs = [] for col_idx in range(X.shape[1]): # 转换为字符串类型避免非文本数据报错 col_data = X[:, col_idx].astype(str) vec = vectorizer.fit_transform(col_data).toarray() X_vecs.append(vec) # 横向拼接所有列的特征 X = np.hstack(X_vecs) tfidfconverter = TfidfTransformer() X = tfidfconverter.fit_transform(X).toarray() print(X.shape) # 输出(10000, 6000),样本数与y一致 # 拆分数据集 X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0)
内容的提问来源于stack exchange,提问作者Pavan Sheelavantar
相关产品推荐
相关产品推荐

