You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GridSearchCV未拟合报错求助:Coursera作业代码问题排查

解决GridSearchCV的NotFittedError错误

问题描述

在Coursera课程作业中,运行代码时第16行触发NotFittedError,错误提示:

NotFittedError: This GridSearchCV instance is not fitted yet. Call 'fit' with appropriate arguments before using this estimator.

原代码如下:

from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import roc_curve, auc

def engagement_model():
    train = pd.read_csv('assets/train.csv')
    train_X = train[train.columns[1:9]]
    train_y = train.iloc[:, 9:]
    test = pd.read_csv('assets/test.csv')
    
    X_train, X_test, y_train, y_test = train_test_split(train_X, train_y)
    class_rf=RandomForestClassifier()
    grid_values = {'n_estimators':[10,100], 'max_depth': [None, 30]}
    grid_clf_auc = GridSearchCV(class_rf, param_grid=grid_values, scoring='roc_auc_score')
    predict_test = grid_clf_auc.predict_proba(test[test.columns[1:9]])
    predict_test = predict_test[:,1]
    
    return pd.series(predict_test, index=[test['id']])
    
engagement_model()

错误原因

  1. 未执行模型拟合:创建GridSearchCV实例后,必须先用训练数据调用fit()方法完成参数调优与模型训练,才能调用预测方法,这是触发错误的核心原因。
  2. 评分参数错误:GridSearchCV的scoring参数应使用sklearn内置的指标名称'roc_auc',而非完整函数名'roc_auc_score'。
  3. 其他细节问题:
    • 遗漏pandas导入语句,导致pd对象未定义;
    • 返回语句中pd.series应为大写开头的pd.Series;
    • train_test_split生成的验证集数据未被使用,可用于模型性能评估。

修正后的代码

from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.ensemble import RandomForestClassifier
import pandas as pd  # 补充pandas导入

def engagement_model():
    train = pd.read_csv('assets/train.csv')
    train_X = train[train.columns[1:9]]
    train_y = train.iloc[:, 9:]
    test = pd.read_csv('assets/test.csv')
    
    X_train, X_val, y_train, y_val = train_test_split(train_X, train_y, random_state=42)
    class_rf = RandomForestClassifier(random_state=42)
    grid_values = {'n_estimators':[10,100], 'max_depth': [None, 30]}
    # 修正scoring参数为内置指标名称
    grid_clf_auc = GridSearchCV(class_rf, param_grid=grid_values, scoring='roc_auc', cv=5)
    # 核心步骤:用训练数据拟合模型
    grid_clf_auc.fit(X_train, y_train.values.ravel())  # ravel()处理标签维度问题
    
    # 可选:用验证集评估模型性能
    print(f"验证集ROC-AUC得分: {grid_clf_auc.score(X_val, y_val.values.ravel())}")
    
    # 对测试集执行预测
    predict_test = grid_clf_auc.predict_proba(test[test.columns[1:9]])[:,1]
    
    # 修正Series写法,索引无需嵌套列表
    return pd.Series(predict_test, index=test['id'])
    
engagement_model()

关键说明

  • grid_clf_auc.fit(X_train, y_train.values.ravel()):这行是解决错误的核心,完成模型的参数调优与训练;ravel()用于将DataFrame格式的标签转为一维数组,符合sklearn的输入要求。
  • 增加random_state参数保证结果可复现,适配作业场景的一致性需求。

内容的提问来源于stack exchange,提问作者Arpit Baheti

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 22:55:16