You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

sklearn Pipeline结合RFECV运行时报特征维度不匹配错误求助

问题根因

维度不匹配报错本质是特征选择流程未与模型训练、预测流程做严格封装:

  • 手动在训练集完成RFECV拟合转换后,预测阶段未将同一拟合好的RFECV应用到测试集做特征筛选,直接将原始维度的测试集输入给仅接受筛选后特征的随机森林,触发维度不匹配
  • GridSearchCV是超参数搜索的包装类,本身没有fit_transform方法,所有特征转换逻辑必须封装在它传入的estimator(即Pipeline)内部,禁止直接对GridSearchCV实例调用转换接口。
正确代码修正方案

所有预处理、特征选择步骤必须全部封装入Pipeline,再将Pipeline作为estimator传入GridSearchCV开展嵌套交叉验证。该模式下,每一轮交叉验证折的特征选择器都会仅在训练折拟合,自动应用到验证折/测试集,完全避免维度错配问题,无需手动调用转换接口。

import numpy as np
from sklearn.datasets import make_classification
from sklearn.pipeline import Pipeline
from sklearn.feature_selection import RFECV
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import GridSearchCV, StratifiedKFold, cross_val_predict
from sklearn.metrics import classification_report, roc_auc_score
import shap
import matplotlib.pyplot as plt

# 构建Pipeline:先执行RFECV特征选择,再输入随机森林分类
pipeline = Pipeline([
    ('feature_selection', RFECV(
        estimator=RandomForestClassifier(random_state=42),
        step=1,
        cv=StratifiedKFold(5, shuffle=True, random_state=42),
        scoring='accuracy'
    )),
    ('classifier', RandomForestClassifier(random_state=42))
])

# 定义超参数搜索空间,参数名必须加*「步骤名__」*前缀匹配Pipeline内对应步骤
param_grid = {
    'feature_selection__min_features_to_select': [5, 10, 14],
    'classifier__n_estimators': [100, 200],
    'classifier__max_depth': [None, 5, 10]
}

# 内层交叉验证:超参数搜索
inner_cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
grid_search = GridSearchCV(
    estimator=pipeline,
    param_grid=param_grid,
    cv=inner_cv,
    scoring='accuracy',
    n_jobs=-1
)

# 外层交叉验证:泛化能力评估
outer_cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
# 示例数据,替换为自有数据集即可,若X为DataFrame可保留列名用于后续特征输出
X, y = make_classification(n_samples=1000, n_features=20, n_informative=12, n_classes=3, random_state=42)

# 直接通过cross_val_predict获取外层交叉验证的预测结果,自动完成全流程转换
y_pred = cross_val_predict(grid_search, X, y, cv=outer_cv, n_jobs=-1, method='predict')
y_pred_proba = cross_val_predict(grid_search, X, y, cv=outer_cv, n_jobs=-1, method='predict_proba')

# 若需手动拆分训练测试集,拟合后直接调用predict即可,无需手动做特征转换
# grid_search.fit(X_train, y_train)
# y_test_pred = grid_search.predict(X_test)
打印RFECV选中的特征列表

嵌套交叉验证每一折选中的特征存在差异,不要拿某一折的特征选择结果作为全量模型的最终特征,若要获取全量数据拟合后最优模型的选中特征,可在全量数据拟合完成后,从最优Pipeline的对应步骤提取特征掩码:

# 在全量数据集上拟合最终最优模型
grid_search.fit(X, y)

# 提取最优Pipeline、RFECV步骤
best_pipe = grid_search.best_estimator_
rfecv_step = best_pipe.named_steps['feature_selection']

# 生成选中特征列表,若X为带列名的DataFrame,将np.arange(X.shape[1])替换为X.columns即可
selected_feature_mask = rfecv_step.support_
selected_features = np.arange(X.shape[1])[selected_feature_mask]

print(f"RFECV最终选中特征数:{rfecv_step.n_features_}")
print(f"RFECV选中特征索引/名称:{selected_features.tolist()}")
后续扩展功能注意事项
  • 多分类指标计算:直接使用外层交叉验证输出的y_pred、y_pred_proba调用sklearn.metrics接口即可,计算多分类AUC、ROC时需指定multi_class='ovr'或multi_class='ovo'参数
  • ROC曲线绘制:直接使用交叉验证返回的预测概率绘制即可,无需手动拆分数据做特征转换
  • SHAP特征重要性计算:不能直接将原始全量特征输入SHAP解释器,需先通过Pipeline完成特征筛选再传入,参考代码如下:
    # 提取最优模型中的随机森林分类器,以及筛选后的特征集
    best_rf = best_pipe.named_steps['classifier']
    X_selected = best_pipe[:-1].transform(X) # 调用Pipeline除最后分类器外的所有步骤做转换
    
    explainer = shap.TreeExplainer(best_rf)
    shap_values = explainer.shap_values(X_selected)
    # 绘图时传入筛选后的特征及对应名称
    shap.summary_plot(shap_values, X_selected, feature_names=selected_features)
    

内容的提问来源于stack exchange,提问作者Slowat_Kela

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 01:39:20