线性建模遇ValueError:输入样本数量不一致问题求助
线性回归代码报错:ValueError: Found input variables with inconsistent numbers of samples: [1, 14]
问题代码
Predictors = pd.DataFrame(['house_age'],['Distance_to_MRT_Station'],['Convenience_Stores']) X=Predictors y=('Price_Per_Sqft') X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.20, random_state=1) lr = LinearRegression() lr.fit(X_train, y_train) LinearRegression()
报错信息
ValueError Traceback (most recent call last) <ipython-input-172-040724a8c0e4> in <module> ----> 1 X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.20, random_state=1) 2 lr = LinearRegression() 3 lr.fit(X_train, y_train) 4 LinearRegression() 2 frames /usr/local/lib/python3.8/dist-packages/sklearn/utils/validation.py in check_consistent_length(*arrays) 330 uniques = np.unique(lengths) 331 if len(uniques) > 1: ---> 332 raise ValueError( 333 "Found input variables with inconsistent numbers of samples: %r" 334 % [int(l) for l in lengths] ValueError: Found input variables with inconsistent numbers of samples: [1, 14]
修复方案
错误根源
- 特征集X创建错误:你用
pd.DataFrame的方式生成的是仅1行3列的空数据框(只有列名,无实际样本数据),样本数为1。 - 目标变量y赋值错误:
y=('Price_Per_Sqft')只是把y定义成了一个字符串,sklearn会把字符串长度(14)当作样本数,导致X和y样本数不匹配(1 vs 14)。
修正代码
假设你的原始数据集已经读取为df(比如通过pd.read_csv加载),正确代码如下:
import pandas as pd from sklearn.model_selection import train_test_split from sklearn.linear_model import LinearRegression # 从原始数据集中提取特征列和目标列 X = df[['house_age', 'Distance_to_MRT_Station', 'Convenience_Stores']] y = df['Price_Per_Sqft'] # 划分训练集和测试集 X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.20, random_state=1) # 训练线性回归模型 lr = LinearRegression() lr.fit(X_train, y_train)
说明
- 必须从已加载的完整数据集中提取特征和目标变量,而非手动创建仅含列名的空数据框。
- 目标变量y要对应数据集中的真实列,不能是字符串字面量。
内容的提问来源于stack exchange,提问作者Hannah Tomlinson
相关产品推荐
相关产品推荐

