Scikit-learn中无法用FeatureUnion结合数值与分类特征求助
Hey there! Let's work through your Pipeline and FeatureUnion issue for predicting Med using Age and Gender—this is a super common roadblock when getting started with scikit-learn's composite estimators, so you're not alone.
First, let's start with a complete, working example that matches your use case. I'll break it down so you can see exactly how data flows through the pipeline, and how to avoid those parameter naming headaches.
Step 1: Build Your Preprocessing + Model Pipeline
Since Age is a numerical feature and Gender is categorical, we'll handle each with separate preprocessing steps, then combine them with FeatureUnion, before feeding into your prediction model.
First, we need a simple helper to select specific columns (since scikit-learn transformers work on arrays, not DataFrame column names by default):
from sklearn.pipeline import Pipeline, FeatureUnion from sklearn.preprocessing import StandardScaler, OneHotEncoder from sklearn.base import BaseEstimator, TransformerMixin from sklearn.linear_model import LogisticRegression from sklearn.model_selection import train_test_split import pandas as pd # Custom transformer to pick specific columns from a DataFrame class ColumnSelector(BaseEstimator, TransformerMixin): def __init__(self, columns): self.columns = columns def fit(self, X, y=None): return self def transform(self, X): return X[self.columns] # Preprocessing for numerical features (Age) num_pipeline = Pipeline([ ('select_num', ColumnSelector(columns=['Age'])), ('scale', StandardScaler()) ]) # Preprocessing for categorical features (Gender) cat_pipeline = Pipeline([ ('select_cat', ColumnSelector(columns=['Gender'])), ('encode', OneHotEncoder(sparse_output=False, drop='first')) ]) # Combine both preprocessing pipelines preprocessor = FeatureUnion([ ('numerical', num_pipeline), ('categorical', cat_pipeline) ]) # Full end-to-end pipeline: preprocessing → model full_pipeline = Pipeline([ ('preprocess', preprocessor), ('model', LogisticRegression()) # Replace with your model of choice ])
Step 2: Split Data and Fit the Pipeline
This is the part where many folks get stuck—you don't need to manually preprocess your train/test splits! The pipeline handles all that automatically:
# Assume your data is stored in a pandas DataFrame called `df` X = df[['Age', 'Gender']] y = df['Med'] # Split your data like normal X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42) # Fit the entire pipeline on training data full_pipeline.fit(X_train, y_train) # Make predictions on test data y_pred = full_pipeline.predict(X_test)
Fixing the model__ Parameter Error
You mentioned switching to model__ prefixes and hitting a new error. Here's the most likely issue and fix:
- The
model__prefix only works if your pipeline's model step is actually namedmodel(like in the example above). If you named it something else (e.g.,classifier,regressor), you need to match that name in the prefix (e.g.,classifier__C). - If you're using
GridSearchCVto tune model parameters, double-check that the parameter name exists for your model. For example,LogisticRegressionusesCfor regularization, butRandomForestClassifierusesn_estimators—mismatched parameter names will throw errors.
Example of valid parameter tuning with model__ prefix:
from sklearn.model_selection import GridSearchCV param_grid = { 'model__C': [0.01, 0.1, 1, 10], # Regularization strength for LogisticRegression 'model__penalty': ['l1', 'l2'] } grid_search = GridSearchCV(full_pipeline, param_grid, cv=5) grid_search.fit(X_train, y_train)
If you can share the exact error message and a snippet of your code where the error occurs, I can help pinpoint the exact issue—but this setup should cover the most common pitfalls.
内容的提问来源于stack exchange,提问作者KubiK888

