机器学习建模报错:输入变量样本数量不一致问题排查
问题:线性回归建模报错ValueError: 输入变量样本数不一致
运行线性回归代码时出现如下报错:
ValueError: Found input variables with inconsistent numbers of samples: [6396, 1599]
报错原因
核心错误是train_test_split的返回值接收顺序完全错误。该函数的正确返回顺序是:X_train, X_test, y_train, y_test,但你写成了X_train, y_train, X_test, y_test,导致:
X_train实际是80%的特征数据(样本数6396)y_train实际是20%的标签数据(样本数1599)
两者样本数不匹配,触发scikit-learn的样本一致性校验报错。
解决方案
修正train_test_split的变量接收顺序,严格对应官方返回的四个值顺序即可。
修正后的完整代码
import pandas as pd import seaborn as sns import matplotlib.pyplot as plt import numpy as np df = pd.read_csv('Armenian Market Car Prices.csv') df['Car Name'] = df['Car Name'].astype('category').cat.codes df = df.join(pd.get_dummies(df.FuelType, dtype=int)) df = df.drop('FuelType', axis=1) df['Region'] = df['Region'].astype('category').cat.codes df['Price'] = df.pop('Price') X = df.drop('Price', axis=1) y = df['Price'] from sklearn.model_selection import train_test_split from sklearn.linear_model import LinearRegression # 修正顺序:X_train, X_test, y_train, y_test X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2) model = LinearRegression() model.fit(X_train, y_train)
内容的提问来源于stack exchange,提问作者Adrian Zambrana
相关产品推荐
相关产品推荐

