You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Sklearn独热编码后如何保留列名?(Kaggle泰坦尼克竞赛场景)

Solution to Keep Feature Names When Merging Encoded Categorical Data with Numeric Data

Hey there! I totally get the frustration of losing column names when combining arrays and DataFrames—especially when you're trying to keep track of features for model interpretability in the Titanic Kaggle competition. Here are two reliable methods to merge your data while preserving feature names, without using pd.get_dummies():

Method 1: Convert Encoded ndarray to DataFrame with Generated Feature Names

Sklearn's OneHotEncoder (version 1.0+) has a handy get_feature_names_out() method that lets you retrieve the exact names of the one-hot encoded columns. Here's how to use it:

from sklearn.preprocessing import OneHotEncoder
import pandas as pd

# Assume X_train_num is your numeric DataFrame, X_train_cat is your categorical DataFrame
# Initialize the encoder (set sparse_output=False to get a dense array)
encoder = OneHotEncoder(sparse_output=False, handle_unknown='ignore')

# Fit and transform the categorical data
X_train_cat_encoded = encoder.fit_transform(X_train_cat)

# Get the feature names for the encoded columns
encoded_feature_names = encoder.get_feature_names_out(input_features=X_train_cat.columns)

# Convert the encoded array to a DataFrame with the correct column names
X_train_cat_df = pd.DataFrame(
    X_train_cat_encoded,
    columns=encoded_feature_names,
    index=X_train_num.index  # Match the index to avoid alignment issues
)

# Merge the numeric and encoded categorical DataFrames
final_df = pd.concat([X_train_num, X_train_cat_df], axis=1)

This works because we explicitly map the encoded array to a DataFrame with meaningful column names (like Sex_female, Embarked_S) before merging, so pd.concat() preserves all column labels.

Method 2: Use ColumnTransformer to Handle Preprocessing in One Step

If you want to streamline your preprocessing pipeline, ColumnTransformer can handle both numeric and categorical variables in one go, and you can easily extract the full list of feature names afterward:

from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder
import pandas as pd
import numpy as np

# Combine your numeric and categorical DataFrames temporarily (just for the transformer)
combined_train = pd.concat([X_train_num, X_train_cat], axis=1)

# Define the preprocessing transformer
preprocessor = ColumnTransformer(
    transformers=[
        # Apply OneHotEncoder to categorical columns
        ('categorical', OneHotEncoder(sparse_output=False, handle_unknown='ignore'), X_train_cat.columns)
    ],
    remainder='passthrough'  # Keep numeric columns as-is
)

# Fit and transform the combined data
processed_array = preprocessor.fit_transform(combined_train)

# Extract all feature names: encoded categorical + original numeric
encoded_cat_names = preprocessor.named_transformers_['categorical'].get_feature_names_out(X_train_cat.columns)
numeric_names = X_train_num.columns
all_features = np.concatenate([encoded_cat_names, numeric_names])

# Convert the processed array back to a DataFrame with full feature names
final_df = pd.DataFrame(processed_array, columns=all_features, index=combined_train.index)

Bonus Tip for Sparse Arrays

If you're working with large datasets and prefer sparse matrices (set sparse_output=True in OneHotEncoder), you can create a sparse DataFrame instead of a dense one to save memory:

X_train_cat_df = pd.DataFrame.sparse.from_spmatrix(
    X_train_cat_encoded,
    columns=encoded_feature_names,
    index=X_train_num.index
)

This way, you keep all your feature names intact while maintaining efficiency.

内容的提问来源于stack exchange,提问作者user2486276

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:33:01