You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修复sklearn线性回归训练时样本数不一致的ValueError报错

线性回归房价预测fit报错排查

问题现象

使用线性回归实现房价预测任务时,调用model.fit()环节触发样本长度不匹配错误,模型训练无法正常完成。

问题复现代码

#importing dependencies
import pandas as pd
import numpy as np
from sklearn.linear_model import LinearRegression
import matplotlib.pyplot as plt
from sklearn.model_selection import train_test_split

#data loading
dataset = pd.read_csv('/content/dataset - train.csv')

#data visualization
plt.xlabel('Area')
plt.ylabel('Price')
plt.scatter(dataset['LotArea'], dataset['SalePrice'], color='red', marker='*')

#splitting data into features and target
X = dataset.drop(['SalePrice'], axis = 1)
Y = dataset['LotArea']

#data splitting into train and test data
X_train, X_test, Y_train, Y_train = train_test_split(X, Y, test_size=0.2, random_state=0)

#training the model
model = LinearRegression()
model.fit(X_train, Y_train)

抛出的错误信息

ValueError                                Traceback (most recent call last)
<ipython-input-31-a42a894194a6> in <module>()
      1 model = LinearRegression()
----> 2 model.fit(X_train, Y_train)

3 frames
/usr/local/lib/python3.7/dist-packages/sklearn/utils/validation.py in check_consistent_length(*arrays)
    332         raise ValueError(
    333             "Found input variables with inconsistent numbers of samples: %r"
--> 334             % [int(l) for l in lengths]
    335         )
    336 

ValueError: Found input variables with inconsistent numbers of samples: [1168, 292]

根因定位

代码里有两处错误,其中第一处直接触发当前报错:

  • 数据集拆分的返回值接收变量写错:train_test_split会按固定顺序返回4个结果:训练集特征、测试集特征、训练集标签、测试集标签,对应接收变量应该依次是X_train, X_test, Y_train, Y_test。你的代码里把第四个返回值错写成了Y_train,导致原本赋值为80%样本量(1168条)的训练集标签Y_train,立刻被覆盖成了仅占20%样本量(292条)的测试集标签。传入model.fit()的特征是1168条、标签是292条,样本数量不匹配直接触发报错。
  • 特征与目标变量的定义不符合任务逻辑:你当前代码把LotArea(房屋面积)设为预测目标Y,把剔除SalePrice(房价)的所有字段设为特征X,相当于要训练模型用其他字段预测房屋面积,和“房价预测”的目标完全相悖,和你之前可视化时“面积为横轴、房价为纵轴”的逻辑也不匹配。

修复方案

根据你的实际任务需求调整代码即可:

场景1:多特征预测房价(用除房价外的所有字段训练模型)

修正特征目标定义、修正拆分语句的变量接收即可:

# 正确定义预测目标为房价SalePrice
X = dataset.drop(['SalePrice'], axis = 1)
Y = dataset['SalePrice']

# 修正返回值接收,不要重复写Y_train
X_train, X_test, Y_train, Y_test = train_test_split(X, Y, test_size=0.2, random_state=0)

model = LinearRegression()
model.fit(X_train, Y_train)

注意:多特征场景下需要提前处理数据集里的非数值类型字段、缺失值,否则后续会触发类型相关报错。

场景2:单特征预测房价(仅用房屋面积预测房价,和你的可视化逻辑一致)

除了修正拆分语句外,还要调整特征X的取值,注意单特征输入需要保持二维结构:

# 特征为房屋面积,目标为房价
X = dataset[['LotArea']]  # 双中括号保证返回二维结构,符合模型输入要求
Y = dataset['SalePrice']

X_train, X_test, Y_train, Y_test = train_test_split(X, Y, test_size=0.2, random_state=0)

model = LinearRegression()
model.fit(X_train, Y_train)

内容的提问来源于stack exchange,提问作者sarahcodebyte

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 05:42:15