Python:使用DataFrame变量进行线性回归的报错解决咨询
解决LinearRegression拟合时的维度错误与Series无reshape属性问题
问题重现
你在DataFrame中有A和B两个变量,执行以下代码时出现错误:
x=df.['A'] # 语法错误:多余的点,正确应为df['A'] y=df.['B'] M = LinearRegression() M.fit(A,B) # 变量名错误:应使用定义好的x、y
报错信息:
Expected 2D array, got 1D array instead:
array=[86. 0. 86. ... 0. 0. 0.].
Reshape your data either using array.reshape(-1, 1) if your data has a single feature or array.reshape(1, -1) if it contains a single sample.
尝试用reshape时又报错:
"'Series' object has no attribute 'reshape'"
原因分析
- 维度要求:sklearn的
LinearRegression.fit()要求特征数据必须是2维数组,但直接从DataFrame取单列得到的是pandas Series(1维结构),不符合要求。 - 方法适配问题:
reshape是numpy数组的方法,Series本身没有这个属性,需要先转换结构再调整维度。
解决办法
以下是几种可行的解决方案:
方案1:双中括号提取列,直接得到2D DataFrame
提取列时使用df[['A']](双中括号),得到的是DataFrame而非Series,天然满足2维要求:
from sklearn.linear_model import LinearRegression # 正确提取2维特征数据 x = df[['A']] y = df['B'] # 目标变量可以是1维Series M = LinearRegression() M.fit(x, y)
方案2:将Series转为numpy数组后reshape
通过.to_numpy()或.values把Series转成numpy数组,再用reshape(-1,1)调整为2维:
x = df['A'].to_numpy().reshape(-1, 1) y = df['B'] M = LinearRegression() M.fit(x, y)
方案3:链式调用Series的values属性与reshape
直接对Series的values属性调用reshape方法:
x = df['A'].values.reshape(-1, 1) y = df['B'] M = LinearRegression() M.fit(x, y)
额外修正点
- 原代码中
df.['A']是语法错误,必须去掉点,写成df['A']; fit(A,B)中的变量名错误,应使用你定义的x和y,否则会报变量未定义错误。
内容的提问来源于stack exchange,提问作者ViB
相关产品推荐
相关产品推荐

