You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在GridSearchCV中将特征集设为超参数并完成对应特征预处理?

解决方案

问题出在你把FeatureSelector嵌套进ColumnTransformer里了——ColumnTransformer是对原始数据集的不同列组分别执行预处理后拼接结果,你的特征选择逻辑根本没办法覆盖全局,自然会保留所有原始列。正确的流程应该是先全局筛选特征,再对筛选后的特征按类型做针对性预处理,具体实现步骤如下:

1. 重构Pipeline结构

把FeatureSelector放在Pipeline的最开头,先完成特征筛选,再把筛选后的特征交给ColumnTransformer做缺失值填充、标准化、独热编码等操作,最后接模型。整个流程是串行的:筛选→预处理→建模。

2. 确保FeatureSelector保留特征名

如果你的FeatureSelector是基于列名筛选的,一定要让它的transform方法返回带列名的DataFrame(而非numpy数组),这样后面的ColumnTransformer才能准确识别数值/类别特征。示例自定义类:

import pandas as pd
from sklearn.base import BaseEstimator, TransformerMixin

class FeatureSelector(BaseEstimator, TransformerMixin):
    def __init__(self, feature_names):
        self.feature_names = feature_names
    
    def fit(self, X, y=None):
        return self
    
    def transform(self, X):
        # 返回筛选后的DataFrame,保留列名
        return X[self.feature_names]

3. 构建完整Pipeline&参数网格

把筛选、预处理、模型串起来,然后在param_grid里指定不同的特征集,以及模型的超参数:

from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.linear_model import LogisticRegression
from sklearn.dummy import DummyClassifier
from sklearn.model_selection import GridSearchCV

# 假设原始数据里的数值/类别基准列(后续筛选的特征从这些里选)
NUM_FEATURES_BASE = ['age', 'income', 'score']
CAT_FEATURES_BASE = ['gender', 'education']

# 定义预处理逻辑:针对筛选后的特征自动匹配对应处理规则
preprocessor = ColumnTransformer(
    transformers=[
        ('num', Pipeline(steps=[
            ('imputer', SimpleImputer(strategy='median')),
            ('scaler', StandardScaler())
        ]), NUM_FEATURES_BASE),  # 写基准列名,筛选后不存在的列会自动被忽略
        ('cat', Pipeline(steps=[
            ('imputer', SimpleImputer(strategy='most_frequent')),
            ('encoder', OneHotEncoder(handle_unknown='ignore'))
        ]), CAT_FEATURES_BASE)
    ])

# 完整Pipeline
pipe = Pipeline([
    ('feature_selector', FeatureSelector(feature_names=[])),  # 初始占位,后续网格搜索替换
    ('preprocessor', preprocessor),
    ('classifier', LogisticRegression())
])

# 参数网格:指定不同特征集+不同模型+模型超参数
param_grid = [
    {
        'feature_selector__feature_names': [
            ['age', 'gender'],
            ['income', 'score', 'education'],
            ['age', 'income', 'gender', 'education']
        ],
        'classifier': [LogisticRegression()],
        'classifier__C': [0.1, 1, 10]
    },
    {
        'feature_selector__feature_names': [
            ['age', 'gender'],
            ['income', 'education']
        ],
        'classifier': [DummyClassifier()],
        'classifier__strategy': ['most_frequent', 'stratified']
    }
]

# 运行网格搜索
grid_search = GridSearchCV(pipe, param_grid, cv=5, scoring='accuracy')
grid_search.fit(X_train, y_train)

关键说明

  • ColumnTransformer里的基准列是原始数据的全量数值/类别列,当FeatureSelector筛选后,不存在的列会自动跳过,不会报错。
  • 这种结构下,每一组网格参数都会先筛选指定特征,再对这些特征做对应预处理,完全匹配你的需求。
  • 如果你是基于索引筛选特征,只需修改FeatureSelector的transform方法,同时让ColumnTransformer用索引指定列组即可。

内容的提问来源于stack exchange,提问作者Amok

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 17:03:18