如何在Scikit-learn网格搜索中集成特征选择器,同时支持无特征选择的配置以评估全部特征的效用?
这个问题确实很常见——当我们想用网格搜索对比特征选择模型和全特征模型时,SequentialFeatureSelector的限制(要求n_features_to_select必须小于总特征数)确实有点棘手。不过Scikit-learn本身就提供了非常优雅的解决方案,用参数网格列表来分别定义两种不同的管道配置就行,我来一步步给你讲清楚:
核心思路
我们需要让网格搜索同时覆盖两种完全不同的管道结构:
- 预处理 →
SequentialFeatureSelector特征选择 → 分类器 - 预处理 → 直接跳过特征选择(用
'passthrough') → 分类器
Scikit-learn的GridSearchCV支持传入参数网格的列表,每个列表元素对应一种管道配置的参数空间,这样就能完美实现“同时评估两种配置”的需求,而且完全符合Scikit-learn的惯用写法。
修改后的完整代码
下面是基于你提供的代码修改后的版本,重点改动在参数网格的定义部分:
import seaborn as sns from sklearn.pipeline import Pipeline from sklearn.model_selection import GridSearchCV from sklearn.linear_model import LogisticRegression from sklearn.feature_selection import SequentialFeatureSelector from sklearn.preprocessing import StandardScaler, OneHotEncoder from sklearn.compose import ColumnTransformer from sklearn.impute import SimpleImputer # Load the Titanic dataset titanic = sns.load_dataset('titanic') # Select features and target features = ['age', 'fare', 'sex'] X = titanic[features] y = titanic['survived'] # Preprocessing pipelines for numeric and categorical features numeric_features = ['age', 'fare'] numeric_transformer = Pipeline(steps=[ ('imputer', SimpleImputer(strategy='constant')), ('scaler', StandardScaler()) ]) categorical_features = ['sex'] categorical_transformer = Pipeline(steps=[ ('imputer', SimpleImputer(strategy='constant')), ('onehot', OneHotEncoder(drop='first')) ]) # Combine preprocessing steps preprocessor = ColumnTransformer(transformers=[ ('num', numeric_transformer, numeric_features), ('cat', categorical_transformer, categorical_features) ]) # Initialize classifier and feature selector clf = LogisticRegression(max_iter=1000, solver='liblinear') sfs = SequentialFeatureSelector(clf, direction='forward') # Create a pipeline that includes preprocessing, feature selection, and classification pipeline = Pipeline(steps=[ ('preprocessor', preprocessor), ('feature_selection', sfs), # 保留这个步骤,但后续可以通过参数替换成passthrough ('classifier', clf) ]) # 关键改动:用参数网格列表,分别定义两种配置的参数空间 param_grid = [ # 配置1:使用SequentialFeatureSelector,指定要选择的特征数 { 'feature_selection': [sfs], 'feature_selection__n_features_to_select': [2], 'classifier__C': [0.1, 1.0, 10.0] }, # 配置2:跳过特征选择,用passthrough直接传递所有特征 { 'feature_selection': ['passthrough'], 'classifier__C': [0.1, 1.0, 10.0] } ] # Create and run the grid search grid_search = GridSearchCV(pipeline, param_grid, cv=5, verbose=1) grid_search.fit(X, y) # Output the best parameters and score print("Best parameters found:", grid_search.best_params_) print("Best cross-validation score:", grid_search.best_score_) # 可选:查看所有候选模型的结果 # import pandas as pd # pd.DataFrame(grid_search.cv_results_)[['param_feature_selection', 'param_classifier__C', 'mean_test_score']]
关键部分解释
参数网格列表的作用:
GridSearchCV会遍历列表中的每个参数网格,分别处理两种配置:- 第一个网格对应使用SFS的情况,会使用你定义的
SequentialFeatureSelector实例,并遍历指定的特征数和正则化参数C; - 第二个网格对应无特征选择的情况,会把
feature_selection步骤替换成'passthrough'(即直接传递预处理后的所有特征),只遍历分类器的C参数。
- 第一个网格对应使用SFS的情况,会使用你定义的
避免参数不匹配错误:
当使用'passthrough'时,我们不需要传递任何feature_selection__xxx参数,这样就不会出现因为SFS限制导致的ValueError,也符合Scikit-learn的参数验证规则。扩展灵活性:
如果想测试更多特征选择的情况(比如同时测试选1个和2个特征),只需要修改第一个网格里的feature_selection__n_features_to_select为[1,2]即可,网格搜索会自动生成所有可能的组合。
运行效果
执行代码后,你会看到网格搜索同时评估了两种配置下的所有参数组合,最终输出最优的参数和得分。比如在泰坦尼克数据集上,全特征模型的得分往往和选2个特征的得分相当,甚至可能略高——这也能帮你判断是否真的需要做特征选择。
备注:内容来源于stack exchange,提问作者Davide Fiocco

