You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scikit-learn中无法用FeatureUnion结合数值与分类特征求助

Hey there! Let's work through your Pipeline and FeatureUnion issue for predicting Med using Age and Gender—this is a super common roadblock when getting started with scikit-learn's composite estimators, so you're not alone.

First, let's start with a complete, working example that matches your use case. I'll break it down so you can see exactly how data flows through the pipeline, and how to avoid those parameter naming headaches.

Step 1: Build Your Preprocessing + Model Pipeline

Since Age is a numerical feature and Gender is categorical, we'll handle each with separate preprocessing steps, then combine them with FeatureUnion, before feeding into your prediction model.

First, we need a simple helper to select specific columns (since scikit-learn transformers work on arrays, not DataFrame column names by default):

from sklearn.pipeline import Pipeline, FeatureUnion
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.base import BaseEstimator, TransformerMixin
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
import pandas as pd

# Custom transformer to pick specific columns from a DataFrame
class ColumnSelector(BaseEstimator, TransformerMixin):
    def __init__(self, columns):
        self.columns = columns
    
    def fit(self, X, y=None):
        return self
    
    def transform(self, X):
        return X[self.columns]

# Preprocessing for numerical features (Age)
num_pipeline = Pipeline([
    ('select_num', ColumnSelector(columns=['Age'])),
    ('scale', StandardScaler())
])

# Preprocessing for categorical features (Gender)
cat_pipeline = Pipeline([
    ('select_cat', ColumnSelector(columns=['Gender'])),
    ('encode', OneHotEncoder(sparse_output=False, drop='first'))
])

# Combine both preprocessing pipelines
preprocessor = FeatureUnion([
    ('numerical', num_pipeline),
    ('categorical', cat_pipeline)
])

# Full end-to-end pipeline: preprocessing → model
full_pipeline = Pipeline([
    ('preprocess', preprocessor),
    ('model', LogisticRegression())  # Replace with your model of choice
])

Step 2: Split Data and Fit the Pipeline

This is the part where many folks get stuck—you don't need to manually preprocess your train/test splits! The pipeline handles all that automatically:

# Assume your data is stored in a pandas DataFrame called `df`
X = df[['Age', 'Gender']]
y = df['Med']

# Split your data like normal
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# Fit the entire pipeline on training data
full_pipeline.fit(X_train, y_train)

# Make predictions on test data
y_pred = full_pipeline.predict(X_test)

Fixing the model__ Parameter Error

You mentioned switching to model__ prefixes and hitting a new error. Here's the most likely issue and fix:

  • The model__ prefix only works if your pipeline's model step is actually named model (like in the example above). If you named it something else (e.g., classifier, regressor), you need to match that name in the prefix (e.g., classifier__C).
  • If you're using GridSearchCV to tune model parameters, double-check that the parameter name exists for your model. For example, LogisticRegression uses C for regularization, but RandomForestClassifier uses n_estimators—mismatched parameter names will throw errors.

Example of valid parameter tuning with model__ prefix:

from sklearn.model_selection import GridSearchCV

param_grid = {
    'model__C': [0.01, 0.1, 1, 10],  # Regularization strength for LogisticRegression
    'model__penalty': ['l1', 'l2']
}

grid_search = GridSearchCV(full_pipeline, param_grid, cv=5)
grid_search.fit(X_train, y_train)

If you can share the exact error message and a snippet of your code where the error occurs, I can help pinpoint the exact issue—but this setup should cover the most common pitfalls.

内容的提问来源于stack exchange,提问作者KubiK888

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:19:46