scikit-learn fit报错:期望二维容器却得到pandas Series的解决咨询
scikit-learn fit方法报错:Expected a 2-dimensional container 解决方法
错误信息
ValueError: Expected a 2-dimensional container but got <class 'pandas.core.series.Series'> instead.
Pass a DataFrame containing a single row (i.e. single sample) or a single column (i.e. single feature) instead.
报错出现在model.fit(X_train, y_train)行,核心原因是scikit-learn模型要求特征数据(X)必须是二维容器(如DataFrame、二维数组),而当前X_train是一维的pandas Series。不需要额外新建DataFrame,以下几种简便方式就能解决问题:
解决方法
1. 从源头获取二维数据
修改切片逻辑,直接得到二维的DataFrame,后续拆分后的X_train自然符合要求:
将原来的x = df.iloc[:,0]改成以下任意一种写法:
# 方式A:用列表索引指定列 x = df.iloc[:, [0]] # 方式B:用切片范围选取列 x = df.iloc[:, 0:1]
注:y可以保持一维Series,LinearRegression对y的一维输入完全兼容。
2. 对已有的Series做转换
如果不想修改前面的切片代码,可在fit前直接转换X_train:
- 用pandas自带的
to_frame()方法:
model.fit(X_train.to_frame(), y_train)
- 借助numpy的
reshape转换为二维数组:
import numpy as np model.fit(np.array(X_train).reshape(-1, 1), y_train)
- 直接调用Series的
values.reshape:
model.fit(X_train.values.reshape(-1, 1), y_train)
修改后的完整代码示例
import pandas as pd import matplotlib.pyplot as plt from sklearn.model_selection import train_test_split from sklearn.linear_model import LinearRegression from sklearn import metrics df = pd.read_excel('Month.xlsx', usecols='W, AJ') # 用列表索引获取二维特征数据 x = df.iloc[:, [0]] y = df.iloc[:, 1] x = x.replace(r'^\s*$', 0, regex=True) y = y.replace(r'^\s*$', 0, regex=True) X_train, X_test, y_train, y_test = train_test_split(x, y, test_size=0.2, random_state=0) model = LinearRegression() model.fit(X_train, y_train) y_pred = model.predict(X_test) plt.scatter(X_test, y_test, color='gray') plt.plot(X_test, y_pred, color='red', linewidth=2) plt.xlabel('Luchtvochtigheid') plt.ylabel('Regenval') plt.show()
内容的提问来源于stack exchange,提问作者kemal ozdogan
相关产品推荐
相关产品推荐

