You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解决train_test_split报错:输入变量样本数不一致[40000,10000]

问题描述

执行代码行X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0)时出现错误:

Error: ValueError: Found input variables with inconsistent numbers of samples: [40000, 10000]

推测是向量化操作后X的数组尺寸变化,导致与y的样本数不匹配。以下是运行输出和代码:

运行输出

(10000, 4)
(10000,)
(40000, 1500)
(10000,)

完整代码

import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.feature_extraction.text import CountVectorizer

# 导入数据集:
dataset = pd.read_excel(r"C:\Users\HPS1RT\Downloads\test\Safety_Prediction.xlsx", nrows=10000)
dataset[["Safety"]] *= 1

# 为X和y变量赋值:
X = dataset.iloc[:, :-1].values
y = dataset.iloc[:, 4].values
print(X.shape)
print(y.shape)
    
# 向量化
vectorizer = CountVectorizer(max_features=1500, min_df=5, max_df=0.7)
X = vectorizer.fit_transform(X.ravel()).toarray()

from sklearn.feature_extraction.text import TfidfTransformer
tfidfconverter = TfidfTransformer()
X = tfidfconverter.fit_transform(X).toarray()
print(X.shape)
print(y.shape)

# 将数据集拆分为随机训练集和测试集:
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0) 
print(X_train)
print(y_train)
问题原因与解决方法

核心问题

你用X.ravel()把原本(10000,4)的二维数组拉平成了40000个元素的一维数组,CountVectorizer会把每个元素当成独立样本处理,导致X变成(40000,1500),和y的10000个样本数量完全不匹配,从而触发错误。

解决思路

根据你的需求选择对应方案:

方案1:合并多列文本为单个样本特征(最常见需求)

如果每行的4个列属于同一样本的不同文本属性,需要先将它们合并成一个字符串,再做向量化:

# 替换原代码中X赋值和向量化的部分
import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.feature_extraction.text import TfidfTransformer

dataset = pd.read_excel(r"C:\Users\HPS1RT\Downloads\test\Safety_Prediction.xlsx", nrows=10000)
dataset[["Safety"]] *= 1

# 将4列文本合并为单个字符串,用空格分隔
X = dataset.iloc[:, :-1].apply(lambda row: ' '.join(row.astype(str)), axis=1).values
y = dataset.iloc[:, 4].values
print(X.shape)  # 输出(10000,)
print(y.shape)  # 输出(10000,)

# 向量化(无需使用ravel)
vectorizer = CountVectorizer(max_features=1500, min_df=5, max_df=0.7)
X = vectorizer.fit_transform(X).toarray()

tfidfconverter = TfidfTransformer()
X = tfidfconverter.fit_transform(X).toarray()
print(X.shape)  # 输出(10000, 1500),与y样本数一致

# 拆分数据集
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0) 

方案2:保留多列独立特征后拼接

如果需要让4个字段分别生成独立的文本特征再合并:

# 替换原代码中向量化的部分
import numpy as np
import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.feature_extraction.text import TfidfTransformer

dataset = pd.read_excel(r"C:\Users\HPS1RT\Downloads\test\Safety_Prediction.xlsx", nrows=10000)
dataset[["Safety"]] *= 1

X = dataset.iloc[:, :-1].values
y = dataset.iloc[:, 4].values
print(X.shape)
print(y.shape)

# 对每一列单独做向量化
vectorizer = CountVectorizer(max_features=1500, min_df=5, max_df=0.7)
X_vecs = []
for col_idx in range(X.shape[1]):
    # 转换为字符串类型避免非文本数据报错
    col_data = X[:, col_idx].astype(str)
    vec = vectorizer.fit_transform(col_data).toarray()
    X_vecs.append(vec)

# 横向拼接所有列的特征
X = np.hstack(X_vecs)

tfidfconverter = TfidfTransformer()
X = tfidfconverter.fit_transform(X).toarray()
print(X.shape)  # 输出(10000, 6000),样本数与y一致

# 拆分数据集
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0) 

内容的提问来源于stack exchange,提问作者Pavan Sheelavantar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 01:24:14