使用sklearn线性回归计算R²时遇特征名匹配错误求助
线性回归模型计算R²时的参数错误解决
问题场景
我尝试用sklearn的LinearRegression构建模型,通过以下步骤计算R²值评估拟合效果:
- 导入CSV数据集
- 筛选目标列(卧室数、浴室数、房价)
- 划分训练集与测试集
- 创建并训练线性回归模型
- 在测试集上生成预测结果
- 计算R²评估模型表现
使用的代码如下:
''' 验证房价与卧室数、浴室数之间的相关性 ''' import pandas as pd from sklearn.model_selection import train_test_split from sklearn.linear_model import LinearRegression df = pd.read_csv('data/American_Housing_Data_20231209.csv') df_interesting_columns = df[['Beds', 'Baths', 'Price']] independent_variables = df_interesting_columns[['Beds', 'Baths']] dependent_variable = df_interesting_columns[['Price']] X_train, X_test, y_train, y_test = train_test_split(independent_variables, dependent_variable, test_size=0.2) model = LinearRegression() model.fit(X_train, y_train) prediction = model.predict(X_test) print(model.score(y_test, prediction))
运行后触发错误:
ValueError: The feature names should match those that were passed during fit.
Feature names unseen at fit time:
- Price
Feature names seen at fit time, yet now missing:- Baths
- Beds
错误原因
你搞反了model.score()的参数顺序:
sklearn中LinearRegression.score()要求第一个参数是特征数据(X),第二个参数是真实标签(y)。你现在传入的y_test是房价标签列(特征名为Price),和模型训练时用的Beds、Baths特征不匹配,因此触发报错。
修正方案
有两种正确的写法:
写法1:直接使用模型的score方法(传入测试集特征和真实标签)
把最后一行代码改成:
print(model.score(X_test, y_test))
写法2:用metrics模块的r2_score函数(传入真实标签和预测值)
先导入metrics模块,再计算:
from sklearn.metrics import r2_score print(r2_score(y_test, prediction))
这两种写法都能正确计算出测试集上的R²值。
内容的提问来源于stack exchange,提问作者Fede
相关产品推荐
相关产品推荐

