绘制散点图触发ValueError:无法转换为Series的问题求助
问题:绘制残差图时出现ValueError错误
我在为感知器制作残差图可视化时,调用plt.scatter()出现错误,错误提示为:ValueError: Unable to coerce to Series, length must be 1: given 50。我推测问题出在X_new、y_new的定义行,疑惑是否需要将n_samples转为数组而非100个独立样本?如果我的理解有误,这个函数实际需要什么参数?
完整报错栈
Traceback (most recent call last) <ipython-input-50-262bfeef3cbc> in <module>() 50 rregr.fit(X_new, y_new) 51 52 preds = rregr.predict(X_new) 53 # order is important! actual - predictions ---> 54 plt.scatter(preds, y-preds) 55 plt.hlines(y=0, xmin=preds.min(), xmax=preds.max()) 56 plt.show() ~/.virtualenvs/dpl-venv/lib/python3.10/site-packages/pandas/core/ops/common.py in <lambda>(self, other) 72 return NotImplemented 73 74 other = item_from_zerodim(other) 75 ---> 76 return method(self, other) ~/.virtualenvs/dpl-venv/lib/python3.10/site-packages/pandas/core/arraylike.py in __sub__(self, other) 192 @unpack_zerodim_and_defer("__sub__") 193 def __sub__(self, other): --> 194 return self._arith_method(other, operator.sub) ~/.virtualenvs/dpl-venv/lib/python3.10/site-packages/pandas/core/frame.py in _arith_method(self, other, op) 7906 ... 8134 msg.format(req_len=len(left.columns), given_len=len(right)) 8135 ) 8136 right = left._constructor_sliced(right, index=left.columns, dtype=dtype) ValueError: Unable to coerce to Series, length must be 1: given 50
代码片段
import pandas as pd import numpy as np import seaborn as sns import matplotlib.pyplot as plt import os from sklearn.linear_model import Perceptron from sklearn.model_selection import train_test_split from sklearn.datasets import make_s_curve, make_regression from sklearn.linear_model import LinearRegression if __name__ == '__main__': path = os.path.join('../data', 'machine-data.csv') df = pd.read_csv(path) ax = sns.heatmap(df.corr(numeric_only=True), annot=True) # histo of output counts = df.loc[:, 'erp'].value_counts() print(counts) sns.barplot(x=counts, y=counts.values) ax = sns.heatmap(df.corr(numeric_only=True).abs(), vmin= 0.65, vmax=1.0, annot=True) in_features = ['myct', 'mmin', 'mmax', 'cach', 'chmin', 'chmax', 'prp'] out_features = ['erp'] X = df.loc[:, in_features] y = df.loc[:, out_features] X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=123857) df.dropna() model = Perceptron() model.fit(X_train, y_train) print('--------------\nTraining Scores') print(model.score(X_train, y_train)) print('--------------\nTesting Scores') print(model.score(X_test, y_test)) X_new, y_new = make_regression(n_samples=100, n_features=1, noise=7) rregr = LinearRegression() rregr.fit(X_new, y_new) preds = rregr.predict(X_new) # order is important! actual - predictions plt.scatter(preds, y-preds) plt.hlines(y=0, xmin=preds.min(), xmax=preds.max()) plt.show()
问题分析与解决方法
核心错误原因
你在计算残差时用了y-preds,但这里的y是之前从CSV读取的DataFrame(形状为[N,1]),而preds是make_regression生成数据的预测结果(形状为[100,]的numpy数组),两者维度和数据完全不匹配,导致pandas在执行减法时抛出长度不匹配的错误。
另外,当前代码逻辑混乱:你前面训练了感知器模型,但后面却用make_regression生成的全新数据训练线性回归并绘制残差图,这和你“为感知器制作残差图”的目标完全无关。
修正方案
方案1:为感知器模型绘制残差图
使用感知器对测试集(或训练集)的预测结果,和对应的真实标签计算残差:
# 用感知器模型对测试集做预测 perceptron_preds = model.predict(X_test) # 计算残差:真实值 - 预测值,将y_test转为一维数组避免维度问题 residuals = y_test.values.ravel() - perceptron_preds # 绘制残差图 plt.scatter(perceptron_preds, residuals) plt.hlines(y=0, xmin=perceptron_preds.min(), xmax=perceptron_preds.max()) plt.xlabel('预测值') plt.ylabel('残差') plt.title('感知器模型残差图') plt.show()
方案2:测试模拟数据的残差图(仅用于验证可视化逻辑)
如果只是想测试make_regression生成数据的残差可视化,要使用该数据对应的真实标签y_new,而非之前的y:
# 修正残差计算 residuals = y_new - preds plt.scatter(preds, residuals) plt.hlines(y=0, xmin=preds.min(), xmax=preds.max()) plt.xlabel('预测值') plt.ylabel('残差') plt.title('模拟回归数据残差图') plt.show()
关于make_regression参数的说明
你的参数设置完全正确:n_samples=100就是生成100个独立样本,函数返回的X_new是形状(100,1)的数组,y_new是形状(100,)的数组,完全符合线性回归的输入要求,不需要修改成“数组而非100个独立样本”。
内容的提问来源于stack exchange,提问作者Cyrxs
相关产品推荐
相关产品推荐

