特征选择交叉验证报AttributeError:dict对象无fit属性
报错根因
报错AttributeError: 'dict' object has no attribute 'fit'的直接原因是代码中rfecv变量被赋值为Python字典,而非实例化的RFECV特征选择对象:你写的键值对格式是Pipeline封装步骤时的配置写法,单独调用特征选择器时,字典本身不具备fit()、transform()这类sklearn估计器方法,调用自然报错。
另外你的函数定义行末尾缺失冒号,修复完上述问题后还会触发语法错误,需要一并补上。
修复步骤
- 将字典形式的rfecv定义改为直接实例化
RFECV类,把原来写在字典里的参数直接传入类构造函数 - 给函数定义行末尾补全冒号
- 建议增加结果存储逻辑,否则嵌套交叉验证跑完后不会输出全局评估结果
- 分类任务建议将内外层的
KFold替换为StratifiedKFold,保证每折类别分布一致,避免结果偏差
修正后的核心代码
# 补全函数定义末尾的冒号 def run_model_with_grid_search(param_grid={},output_plt_file = 'plt.png',model_name=RandomForestClassifier(),X_train=full_X_train,y_train=full_y_train,model_id='random_forest'): cv_outer = StratifiedKFold(n_splits=5,shuffle=True,random_state=1) outer_test_acc = [] for train_ix,test_ix in cv_outer.split(X_train, y_train): split_x_train, split_x_test = X_train[train_ix,:],X_train[test_ix,:] split_y_train, split_y_test = y_train[train_ix],y_train[test_ix] cv_inner = StratifiedKFold(n_splits=3,shuffle=True,random_state=1) model = model_name # 替换字典为RFECV实例化 rfecv = RFECV( estimator=model, step=1, cv=5, scoring='accuracy', verbose=50 ) rfecv.fit(split_x_train,split_y_train) print(f"当前折选中特征数:{rfecv.n_features_}") X_selected_train = rfecv.transform(split_x_train) X_selected_test = rfecv.transform(split_x_test) search = GridSearchCV(model,param_grid=param_grid,scoring='roc_auc',cv=cv_inner,refit=True) result = search.fit(X_selected_train,split_y_train) best_model = result.best_estimator_ y_pred_test = best_model.predict(X_selected_test) accuracy_test = metrics.accuracy_score(split_y_test, y_pred_test) outer_test_acc.append(accuracy_test) print(f"当前折测试集准确率:{accuracy_test:.3f}") print(f"嵌套交叉验证平均测试集准确率:{np.mean(outer_test_acc):.3f} ± {np.std(outer_test_acc):.3f}") return
额外优化提示
- 代码中存在多处重复导入(比如多次导入
SelectKBest、StratifiedKFold、Pipeline),不影响运行但可以清理冗余 - 如果后续要把RFECV和模型封装到Pipeline中统一做网格搜索,参数网格里的参数名需要遵循
<步骤名>__<参数名>的格式,否则会报参数不存在的错误 - 当前参数网格里的
min_samples_leaf是随机森林的原生参数,直接传入GridSearchCV可以正常识别运行
内容的提问来源于stack exchange,提问作者Slowat_Kela
相关产品推荐
相关产品推荐

