You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用sklearn线性回归计算R²时遇特征名匹配错误求助

线性回归模型计算R²时的参数错误解决

问题场景

我尝试用sklearn的LinearRegression构建模型,通过以下步骤计算R²值评估拟合效果:

  • 导入CSV数据集
  • 筛选目标列(卧室数、浴室数、房价)
  • 划分训练集与测试集
  • 创建并训练线性回归模型
  • 在测试集上生成预测结果
  • 计算R²评估模型表现

使用的代码如下:

''' 验证房价与卧室数、浴室数之间的相关性 '''

import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression

df = pd.read_csv('data/American_Housing_Data_20231209.csv')

df_interesting_columns = df[['Beds', 'Baths', 'Price']]

independent_variables = df_interesting_columns[['Beds', 'Baths']]
dependent_variable = df_interesting_columns[['Price']]

X_train, X_test, y_train, y_test = train_test_split(independent_variables, dependent_variable, test_size=0.2)

model = LinearRegression()
model.fit(X_train, y_train)

prediction = model.predict(X_test)

print(model.score(y_test, prediction))

运行后触发错误:

ValueError: The feature names should match those that were passed during fit.
Feature names unseen at fit time:

  • Price
    Feature names seen at fit time, yet now missing:
  • Baths
  • Beds

错误原因

你搞反了model.score()的参数顺序:
sklearn中LinearRegression.score()要求第一个参数是特征数据(X),第二个参数是真实标签(y)。你现在传入的y_test是房价标签列(特征名为Price),和模型训练时用的Beds、Baths特征不匹配,因此触发报错。

修正方案

有两种正确的写法:

写法1:直接使用模型的score方法(传入测试集特征和真实标签)

把最后一行代码改成:

print(model.score(X_test, y_test))

写法2:用metrics模块的r2_score函数(传入真实标签和预测值)

先导入metrics模块,再计算:

from sklearn.metrics import r2_score
print(r2_score(y_test, prediction))

这两种写法都能正确计算出测试集上的R²值。


内容的提问来源于stack exchange,提问作者Fede

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 20:17:22