使用HistGradientBoostingClassifier进行RFECV时遇ValueError问题求助
解决HistGradientBoostingClassifier结合RFECV的ValueError问题
问题原因
你提前拟合了HistGradientBoostingClassifier实例并传给RFECV,但RFECV在交叉验证流程中,会基于传入的estimator模板,为每个交叉验证折创建全新的未拟合实例。这些新实例没有经过拟合,自然不存在feature_importances_属性,从而触发报错。
修正步骤
- 不要提前调用
fit()方法,直接将未拟合的HistGradientBoostingClassifier实例传给RFECV,让RFECV在交叉验证过程中自行完成每个折的拟合操作。 - 修复
X_train.columns的错误:make_classification生成的X_train是numpy数组,没有columns属性,需改用索引标记选中的特征,或者将数组转为DataFrame。
修正后的完整代码
from sklearn.ensemble import HistGradientBoostingClassifier from sklearn.feature_selection import RFECV from sklearn.model_selection import RepeatedStratifiedKFold from sklearn.datasets import make_classification import pandas as pd # 生成数据集 X_train, y_train = make_classification(n_samples=1000, n_features=20, n_informative=10, n_redundant=5, random_state=42) # 转为DataFrame方便用列名展示 X_train = pd.DataFrame(X_train, columns=[f"feature_{i+1}" for i in range(X_train.shape[1])]) # 直接传入未拟合的分类器实例 estimator = HistGradientBoostingClassifier() # 初始化RFECV rfecv = RFECV(estimator=estimator, step=1, cv=RepeatedStratifiedKFold(n_splits=5, n_repeats=1), scoring='roc_auc') # 拟合数据 rfecv.fit(X_train, y_train) # 打印选中的特征和排名 print("Selected Features: ", X_train.columns[rfecv.support_]) print("Feature Rankings: ", rfecv.ranking_)
内容的提问来源于stack exchange,提问作者learner
相关产品推荐
相关产品推荐

