You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

机器学习建模报错:输入变量样本数量不一致问题排查

问题:线性回归建模报错ValueError: 输入变量样本数不一致

运行线性回归代码时出现如下报错:

ValueError: Found input variables with inconsistent numbers of samples: [6396, 1599]

报错原因

核心错误是train_test_split的返回值接收顺序完全错误。该函数的正确返回顺序是:X_train, X_test, y_train, y_test,但你写成了X_train, y_train, X_test, y_test,导致:

  • X_train实际是80%的特征数据(样本数6396)
  • y_train实际是20%的标签数据(样本数1599)
    两者样本数不匹配,触发scikit-learn的样本一致性校验报错。

解决方案

修正train_test_split的变量接收顺序,严格对应官方返回的四个值顺序即可。

修正后的完整代码

import pandas as pd
import seaborn as sns
import matplotlib.pyplot as plt
import numpy as np

df = pd.read_csv('Armenian Market Car Prices.csv')

df['Car Name'] = df['Car Name'].astype('category').cat.codes

df = df.join(pd.get_dummies(df.FuelType, dtype=int))
df = df.drop('FuelType', axis=1)

df['Region'] = df['Region'].astype('category').cat.codes

df['Price'] = df.pop('Price')

X = df.drop('Price', axis=1)
y = df['Price']

from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression

# 修正顺序:X_train, X_test, y_train, y_test
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2)
model = LinearRegression()

model.fit(X_train, y_train)

内容的提问来源于stack exchange,提问作者Adrian Zambrana

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 11:13:21