使用sklearn train_test_split触发ValueError:期望2D数组却得到1D数组
错误原因及解决方案
核心问题分析
你的代码里有三个关键错误导致了这个ValueError:
- train_test_split返回值顺序完全错误:正确的返回顺序是
x_train, x_test, y_train, y_test,你写反后,x_train实际拿到的是标签集y的训练部分(一维Series),直接导致后续标准化时输入了一维数据。 - StandardScaler.fit()参数错误:你传入了
x_train.shape(一个元组),而不是特征数据本身,完全不符合fit方法的要求。 - 多余的reshape操作:正确的特征集
x_train是二维结构(DataFrame),不需要执行reshape(-1,1),这个操作反而会破坏数据结构。
修正后的完整代码
import pandas as pd from sklearn.impute import SimpleImputer from sklearn.preprocessing import MinMaxScaler import seaborn as sb from sklearn.preprocessing import StandardScaler from sklearn.model_selection import train_test_split df = sb.load_dataset('titanic') df2 = df[['survived','pclass','age','parch']] df3 = df2.fillna(df2.mean()) x = df3.drop('survived', axis=1) y = df3['survived'] # 修正train_test_split的返回顺序 x_train, x_test, y_train, y_test = train_test_split(x, y, test_size=0.2, random_state=51) print('x_train shape:', x_train.shape) # 正常输出为(712, 3),是二维结构 sc = StandardScaler() # 用特征数据x_train训练scaler,而非shape参数 sc.fit(x_train) # 直接对二维的x_train和x_test执行标准化 x_train_sc = sc.transform(x_train) x_test_sc = sc.transform(x_test) print(x_train_sc)
额外简化技巧
可以用fit_transform一步完成scaler的训练和数据转换,简化代码:
sc = StandardScaler() x_train_sc = sc.fit_transform(x_train) x_test_sc = sc.transform(x_test)
内容的提问来源于stack exchange,提问作者Dev King
相关产品推荐
相关产品推荐

