You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scikit-learn检测过拟合时遇AttributeError错误,求排查方案

问题排查:LinearRegression无scorer属性错误

我参考scikit-learn交叉验证检测过拟合的代码实现时,运行出现AttributeError错误,代码及错误信息如下:

原代码

lin_regressor = LinearRegression()

poly = PolynomialFeatures(2)

X_transform = poly.fit_transform(x_train)

linear_regg=lin_regressor.fit(X_transform,y_train) 

from sklearn.metrics import SCORERS
from sklearn.model_selection import KFold

scorer = SCORERS['r2']

cv = KFold(n_splits=5, random_state=0,shuffle=True)

train_scores, test_scores = [], []

for train, test in cv.split(X_normalized):
    X_transform2 = poly.fit_transform(X_normalized)
    OL = lin_regressor.fit(X_transform2[train], y_for_normalized.iloc[train])
    tr_21 = OL.scorer(X_transform2[train], y_for_normalized.iloc[train])
    ts_21 = OL.scorer(X_transform2[test], y_for_normalized.iloc[test])
    print("Train score:", tr_21)  
    print("Test score:", ts_21)  

    train_scores.append(tr_21)
    test_scores.append(ts_21)


print("The Mean for Train scores is:", (np.mean(train_scores)))

print("The Mean for Test scores is:", (np.mean(test_scores)))

错误信息

AttributeError                            Traceback (most recent call last)
/var/folders/mm/r4gnnwl948zclfyx12w803040000gn/T/ipykernel_13098/1734683602.py in <module>
     12     # [] for X_transform2, .iloc[] for y_for_normalized
     13     OL = lin_regressor.fit(X_transform2[train], y_for_normalized.iloc[train])
---> 14     tr_21 = OL.scorer(X_transform2[train], y_for_normalized.iloc[train])
     15     ts_21 = OL.scorer(X_transform2[test], y_for_normalized.iloc[test])
     16     print("Train score:", tr_21)  # from documentation .score returns r^2

AttributeError: 'LinearRegression' object has no attribute 'scorer'

错误原因

LinearRegression模型本身没有scorer属性,代码混淆了模型自带的score()方法和从SCORERS中导入的评分器对象,原代码中调用OL.scorer()是错误用法。

修正方案

有两种正确的评分方式,同时优化代码中的冗余操作:

修正后的代码

import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.preprocessing import PolynomialFeatures
from sklearn.metrics import SCORERS
from sklearn.model_selection import KFold

lin_regressor = LinearRegression()
poly = PolynomialFeatures(2)

# 先拟合多项式特征,避免循环内重复拟合
poly.fit(X_normalized)

scorer = SCORERS['r2']
cv = KFold(n_splits=5, random_state=0, shuffle=True)

train_scores, test_scores = [], []

for train_idx, test_idx in cv.split(X_normalized):
    # 仅做特征转换,不再重复拟合
    X_transform = poly.transform(X_normalized)
    # 训练模型
    OL = lin_regressor.fit(X_transform[train_idx], y_for_normalized.iloc[train_idx])
    
    # 方式1:使用模型自带的score方法(LinearRegression默认返回R²评分)
    tr_score = OL.score(X_transform[train_idx], y_for_normalized.iloc[train_idx])
    ts_score = OL.score(X_transform[test_idx], y_for_normalized.iloc[test_idx])
    
    # 方式2:使用导入的scorer对象(与方式1结果一致)
    # tr_score = scorer(OL, X_transform[train_idx], y_for_normalized.iloc[train_idx])
    # ts_score = scorer(OL, X_transform[test_idx], y_for_normalized.iloc[test_idx])
    
    print(f"训练集评分: {tr_score}")  
    print(f"测试集评分: {ts_score}")  

    train_scores.append(tr_score)
    test_scores.append(ts_score)

print("训练集评分均值:", np.mean(train_scores))
print("测试集评分均值:", np.mean(test_scores))

额外优化点

  • 避免在循环内重复调用poly.fit_transform():先对特征集做一次fit,后续仅做transform,防止数据泄露和冗余计算
  • 变量命名更清晰:将循环中的train/test改为train_idx/test_idx,避免与模型变量混淆

内容的提问来源于stack exchange,提问作者JZ0

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 21:55:18