You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

线性回归模型报错:期望二维数组、样本数不一致问题求解

线性回归模型报错解决与解释

一、样本数不一致错误(Found input variables with inconsistent numbers of samples)

问题原因

调用train_test_split时变量赋值顺序错误。sklearn的train_test_split返回顺序固定为:训练特征(X_train)、测试特征(X_test)、训练标签(y_train)、测试标签(y_test),你之前写反了后两个变量,导致y_train实际接收的是X的测试集,y_test接收的是y的测试集,两者样本数自然不匹配。

解决代码

纠正赋值顺序:

from sklearn.model_selection import train_test_split
# 正确顺序:X_train, X_test, y_train, y_test
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size = 0.2)

二、预测时数组维度错误(Expected 2D array, got 1D array instead)

问题原因

sklearn所有模型的predict方法要求输入必须是二维数组,形状为[样本数量, 特征数量]。你传入的[1,10,4]是一维数组(形状为[3]),模型无法识别这是一个包含3个特征的单一样本。

解决方法

把输入转换成二维数组,两种常用方式:

  1. 直接使用双层列表:
LR.predict([[1, 10, 4]])
  1. 用numpy的reshape方法转换:
import numpy as np
sample = np.array([1, 10, 4]).reshape(1, -1)  # 1表示1个样本,-1自动匹配特征数
LR.predict(sample)

额外优化建议:LabelEncoder的正确用法

你当前循环中对两个特征使用同一个LabelEncoder实例,会导致第二个特征的编码基于第一个特征的类别,容易出现编码混淆(比如两个特征有相同名称的类别时,编码错误)。建议每个特征使用独立的编码器:

from sklearn.preprocessing import LabelEncoder

cols = ['nama_pasar','komoditas']
encoder_dict = {}

for col in cols:
    le = LabelEncoder()
    df_test[col] = le.fit_transform(df_test[col])
    encoder_dict[col] = le  # 保存编码器,方便后续逆变换
    print(le.classes_)

内容的提问来源于stack exchange,提问作者Skidud

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 06:50:15