You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python sklearn多项式回归报Expected 2D array错误求助

报错原因
  • sklearn中所有用于模型训练的特征矩阵都要求是二维结构,维度格式为(样本数, 特征数)。你当前通过x=data["level_id"].astype(int)获取到的是一维的Pandas Series对象,维度为(样本数,),不符合输入要求,因此触发维度校验错误。
  • 如果你之前尝试reshape仍失败,大概率是直接对Series对象调用了reshape方法,Pandas 1.0+版本中Series的reshape方法已被废弃,需要先转成Numpy数组再执行维度调整。
修复方案

最简单的改法是在取特征列时直接用双层方括号,拿到的默认是二维的DataFrame结构,不需要额外reshape:
把原代码中取特征的部分:

x=data["level_id"].astype(int)
y=data["Salary"].astype(int)

修改为:

# 双层方括号取列得到二维DataFrame,符合sklearn输入要求
x = data[["level_id"]].astype(int)
y = data["Salary"].astype(int)

如果你习惯用reshape操作,也可以用如下写法:

x = data["level_id"].astype(int).values.reshape(-1, 1)
y = data["Salary"].astype(int)

额外优化建议

你的代码中多次调用了poly.fit_transform(),重复fit会导致多项式生成规则反复更新,建议调整为只fit一次,后续统一用transform,避免出现数据不一致问题。

完整可运行代码

# importing libraries
import numpy as np
import pandas as pd
import matplotlib
matplotlib.use('TkAgg')
import matplotlib.pyplot as plt
from sklearn.preprocessing import PolynomialFeatures
from sklearn.linear_model import LinearRegression

# read data
dataset = pd.read_csv("Salary_Levels.csv")
data = pd.DataFrame(dataset)

# 修正特征维度为二维
x = data[["level_id"]].astype(int)
y = data["Salary"].astype(int)

# 多项式特征转换与模型训练
poly = PolynomialFeatures(degree=4)
x_poly = poly.fit_transform(x)
pilreg = LinearRegression()
pilreg.fit(x_poly, y)
# 预测level=10对应的薪资
print(pilreg.predict(poly.transform([[10]])))

# 绘图
plt.scatter(x, y, color='r', s=5)
plt.plot(x, pilreg.predict(x_poly), color='blue')
plt.show()

内容的提问来源于stack exchange,提问作者David

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 14:54:07