You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Scikit-learn网格搜索中集成特征选择器,同时支持无特征选择的配置以评估全部特征的效用?

如何在Scikit-learn网格搜索中集成特征选择器,同时支持无特征选择的配置以评估全部特征的效用?

这个问题确实很常见——当我们想用网格搜索对比特征选择模型和全特征模型时,SequentialFeatureSelector的限制(要求n_features_to_select必须小于总特征数)确实有点棘手。不过Scikit-learn本身就提供了非常优雅的解决方案,用参数网格列表来分别定义两种不同的管道配置就行,我来一步步给你讲清楚:

核心思路

我们需要让网格搜索同时覆盖两种完全不同的管道结构:

  1. 预处理 → SequentialFeatureSelector 特征选择 → 分类器
  2. 预处理 → 直接跳过特征选择(用'passthrough') → 分类器

Scikit-learn的GridSearchCV支持传入参数网格的列表,每个列表元素对应一种管道配置的参数空间,这样就能完美实现“同时评估两种配置”的需求,而且完全符合Scikit-learn的惯用写法。

修改后的完整代码

下面是基于你提供的代码修改后的版本,重点改动在参数网格的定义部分:

import seaborn as sns
from sklearn.pipeline import Pipeline
from sklearn.model_selection import GridSearchCV
from sklearn.linear_model import LogisticRegression
from sklearn.feature_selection import SequentialFeatureSelector
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer

# Load the Titanic dataset
titanic = sns.load_dataset('titanic')

# Select features and target
features = ['age', 'fare', 'sex']
X = titanic[features]
y = titanic['survived']

# Preprocessing pipelines for numeric and categorical features
numeric_features = ['age', 'fare']
numeric_transformer = Pipeline(steps=[
    ('imputer', SimpleImputer(strategy='constant')),
    ('scaler', StandardScaler())
])

categorical_features = ['sex']
categorical_transformer = Pipeline(steps=[
    ('imputer', SimpleImputer(strategy='constant')),
    ('onehot', OneHotEncoder(drop='first'))
])

# Combine preprocessing steps
preprocessor = ColumnTransformer(transformers=[
    ('num', numeric_transformer, numeric_features),
    ('cat', categorical_transformer, categorical_features)
])

# Initialize classifier and feature selector
clf = LogisticRegression(max_iter=1000, solver='liblinear')
sfs = SequentialFeatureSelector(clf, direction='forward')

# Create a pipeline that includes preprocessing, feature selection, and classification
pipeline = Pipeline(steps=[
    ('preprocessor', preprocessor),
    ('feature_selection', sfs),  # 保留这个步骤,但后续可以通过参数替换成passthrough
    ('classifier', clf)
])

# 关键改动:用参数网格列表,分别定义两种配置的参数空间
param_grid = [
    # 配置1:使用SequentialFeatureSelector,指定要选择的特征数
    {
        'feature_selection': [sfs],
        'feature_selection__n_features_to_select': [2],
        'classifier__C': [0.1, 1.0, 10.0]
    },
    # 配置2:跳过特征选择,用passthrough直接传递所有特征
    {
        'feature_selection': ['passthrough'],
        'classifier__C': [0.1, 1.0, 10.0]
    }
]

# Create and run the grid search
grid_search = GridSearchCV(pipeline, param_grid, cv=5, verbose=1)
grid_search.fit(X, y)

# Output the best parameters and score
print("Best parameters found:", grid_search.best_params_)
print("Best cross-validation score:", grid_search.best_score_)

# 可选:查看所有候选模型的结果
# import pandas as pd
# pd.DataFrame(grid_search.cv_results_)[['param_feature_selection', 'param_classifier__C', 'mean_test_score']]

关键部分解释

  1. 参数网格列表的作用:
    GridSearchCV会遍历列表中的每个参数网格,分别处理两种配置:

    • 第一个网格对应使用SFS的情况,会使用你定义的SequentialFeatureSelector实例,并遍历指定的特征数和正则化参数C;
    • 第二个网格对应无特征选择的情况,会把feature_selection步骤替换成'passthrough'(即直接传递预处理后的所有特征),只遍历分类器的C参数。
  2. 避免参数不匹配错误:
    当使用'passthrough'时,我们不需要传递任何feature_selection__xxx参数,这样就不会出现因为SFS限制导致的ValueError,也符合Scikit-learn的参数验证规则。

  3. 扩展灵活性:
    如果想测试更多特征选择的情况(比如同时测试选1个和2个特征),只需要修改第一个网格里的feature_selection__n_features_to_select为[1,2]即可,网格搜索会自动生成所有可能的组合。

运行效果

执行代码后,你会看到网格搜索同时评估了两种配置下的所有参数组合,最终输出最优的参数和得分。比如在泰坦尼克数据集上,全特征模型的得分往往和选2个特征的得分相当,甚至可能略高——这也能帮你判断是否真的需要做特征选择。

备注:内容来源于stack exchange,提问作者Davide Fiocco

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 13:00:28