You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

绘制散点图触发ValueError:无法转换为Series的问题求助

问题:绘制残差图时出现ValueError错误

我在为感知器制作残差图可视化时,调用plt.scatter()出现错误,错误提示为:ValueError: Unable to coerce to Series, length must be 1: given 50。我推测问题出在X_new、y_new的定义行,疑惑是否需要将n_samples转为数组而非100个独立样本?如果我的理解有误,这个函数实际需要什么参数?

完整报错栈

Traceback (most recent call last)
<ipython-input-50-262bfeef3cbc> in <module>()
     50     rregr.fit(X_new, y_new)
     51 
     52     preds = rregr.predict(X_new)
     53     # order is important!  actual - predictions
---&gt; 54     plt.scatter(preds, y-preds)
     55     plt.hlines(y=0, xmin=preds.min(), xmax=preds.max())
     56     plt.show()

~/.virtualenvs/dpl-venv/lib/python3.10/site-packages/pandas/core/ops/common.py in <lambda>(self, other)
     72                     return NotImplemented
     73 
     74         other = item_from_zerodim(other)
     75 
---&gt; 76         return method(self, other)

~/.virtualenvs/dpl-venv/lib/python3.10/site-packages/pandas/core/arraylike.py in __sub__(self, other)
    192     @unpack_zerodim_and_defer("__sub__")
    193     def __sub__(self, other):
--&gt; 194         return self._arith_method(other, operator.sub)

~/.virtualenvs/dpl-venv/lib/python3.10/site-packages/pandas/core/frame.py in _arith_method(self, other, op)
   7906 
...
   8134                         msg.format(req_len=len(left.columns), given_len=len(right))
   8135                     )
   8136                 right = left._constructor_sliced(right, index=left.columns, dtype=dtype)

ValueError: Unable to coerce to Series, length must be 1: given 50

代码片段

import pandas as pd
import numpy as np
import seaborn as sns
import matplotlib.pyplot as plt
import os
from sklearn.linear_model import Perceptron
from sklearn.model_selection import train_test_split
from sklearn.datasets import make_s_curve, make_regression
from sklearn.linear_model import LinearRegression

if __name__ == '__main__':

    path = os.path.join('../data', 'machine-data.csv')
    df = pd.read_csv(path)

    ax = sns.heatmap(df.corr(numeric_only=True), annot=True)

    # histo of output
    counts = df.loc[:, 'erp'].value_counts()
    print(counts)
    sns.barplot(x=counts, y=counts.values)

    ax = sns.heatmap(df.corr(numeric_only=True).abs(), vmin= 0.65, vmax=1.0, annot=True)

    in_features = ['myct', 'mmin', 'mmax', 'cach', 'chmin', 'chmax', 'prp']
    out_features = ['erp']

    X = df.loc[:, in_features]
    y = df.loc[:, out_features]

    X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=123857)
    df.dropna()

    model = Perceptron()

    model.fit(X_train, y_train)

    print('--------------\nTraining Scores')
    print(model.score(X_train, y_train))
    print('--------------\nTesting Scores')
    print(model.score(X_test, y_test))

    X_new, y_new = make_regression(n_samples=100, n_features=1, noise=7)

    rregr = LinearRegression()
    rregr.fit(X_new, y_new)

    preds = rregr.predict(X_new)
    # order is important!  actual - predictions
    plt.scatter(preds, y-preds)
    plt.hlines(y=0, xmin=preds.min(), xmax=preds.max())
    plt.show()

问题分析与解决方法

核心错误原因

你在计算残差时用了y-preds,但这里的y是之前从CSV读取的DataFrame(形状为[N,1]),而preds是make_regression生成数据的预测结果(形状为[100,]的numpy数组),两者维度和数据完全不匹配,导致pandas在执行减法时抛出长度不匹配的错误。

另外,当前代码逻辑混乱:你前面训练了感知器模型,但后面却用make_regression生成的全新数据训练线性回归并绘制残差图,这和你“为感知器制作残差图”的目标完全无关。

修正方案

方案1:为感知器模型绘制残差图

使用感知器对测试集(或训练集)的预测结果,和对应的真实标签计算残差:

# 用感知器模型对测试集做预测
perceptron_preds = model.predict(X_test)
# 计算残差:真实值 - 预测值,将y_test转为一维数组避免维度问题
residuals = y_test.values.ravel() - perceptron_preds
# 绘制残差图
plt.scatter(perceptron_preds, residuals)
plt.hlines(y=0, xmin=perceptron_preds.min(), xmax=perceptron_preds.max())
plt.xlabel('预测值')
plt.ylabel('残差')
plt.title('感知器模型残差图')
plt.show()

方案2:测试模拟数据的残差图(仅用于验证可视化逻辑)

如果只是想测试make_regression生成数据的残差可视化,要使用该数据对应的真实标签y_new,而非之前的y:

# 修正残差计算
residuals = y_new - preds
plt.scatter(preds, residuals)
plt.hlines(y=0, xmin=preds.min(), xmax=preds.max())
plt.xlabel('预测值')
plt.ylabel('残差')
plt.title('模拟回归数据残差图')
plt.show()

关于make_regression参数的说明

你的参数设置完全正确:n_samples=100就是生成100个独立样本,函数返回的X_new是形状(100,1)的数组,y_new是形状(100,)的数组,完全符合线性回归的输入要求,不需要修改成“数组而非100个独立样本”。

内容的提问来源于stack exchange,提问作者Cyrxs

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 01:52:17