在Sklearn独热编码后如何保留列名?(Kaggle泰坦尼克竞赛场景)
Solution to Keep Feature Names When Merging Encoded Categorical Data with Numeric Data
Hey there! I totally get the frustration of losing column names when combining arrays and DataFrames—especially when you're trying to keep track of features for model interpretability in the Titanic Kaggle competition. Here are two reliable methods to merge your data while preserving feature names, without using pd.get_dummies():
Method 1: Convert Encoded ndarray to DataFrame with Generated Feature Names
Sklearn's OneHotEncoder (version 1.0+) has a handy get_feature_names_out() method that lets you retrieve the exact names of the one-hot encoded columns. Here's how to use it:
from sklearn.preprocessing import OneHotEncoder import pandas as pd # Assume X_train_num is your numeric DataFrame, X_train_cat is your categorical DataFrame # Initialize the encoder (set sparse_output=False to get a dense array) encoder = OneHotEncoder(sparse_output=False, handle_unknown='ignore') # Fit and transform the categorical data X_train_cat_encoded = encoder.fit_transform(X_train_cat) # Get the feature names for the encoded columns encoded_feature_names = encoder.get_feature_names_out(input_features=X_train_cat.columns) # Convert the encoded array to a DataFrame with the correct column names X_train_cat_df = pd.DataFrame( X_train_cat_encoded, columns=encoded_feature_names, index=X_train_num.index # Match the index to avoid alignment issues ) # Merge the numeric and encoded categorical DataFrames final_df = pd.concat([X_train_num, X_train_cat_df], axis=1)
This works because we explicitly map the encoded array to a DataFrame with meaningful column names (like Sex_female, Embarked_S) before merging, so pd.concat() preserves all column labels.
Method 2: Use ColumnTransformer to Handle Preprocessing in One Step
If you want to streamline your preprocessing pipeline, ColumnTransformer can handle both numeric and categorical variables in one go, and you can easily extract the full list of feature names afterward:
from sklearn.compose import ColumnTransformer from sklearn.preprocessing import OneHotEncoder import pandas as pd import numpy as np # Combine your numeric and categorical DataFrames temporarily (just for the transformer) combined_train = pd.concat([X_train_num, X_train_cat], axis=1) # Define the preprocessing transformer preprocessor = ColumnTransformer( transformers=[ # Apply OneHotEncoder to categorical columns ('categorical', OneHotEncoder(sparse_output=False, handle_unknown='ignore'), X_train_cat.columns) ], remainder='passthrough' # Keep numeric columns as-is ) # Fit and transform the combined data processed_array = preprocessor.fit_transform(combined_train) # Extract all feature names: encoded categorical + original numeric encoded_cat_names = preprocessor.named_transformers_['categorical'].get_feature_names_out(X_train_cat.columns) numeric_names = X_train_num.columns all_features = np.concatenate([encoded_cat_names, numeric_names]) # Convert the processed array back to a DataFrame with full feature names final_df = pd.DataFrame(processed_array, columns=all_features, index=combined_train.index)
Bonus Tip for Sparse Arrays
If you're working with large datasets and prefer sparse matrices (set sparse_output=True in OneHotEncoder), you can create a sparse DataFrame instead of a dense one to save memory:
X_train_cat_df = pd.DataFrame.sparse.from_spmatrix( X_train_cat_encoded, columns=encoded_feature_names, index=X_train_num.index )
This way, you keep all your feature names intact while maintaining efficiency.
内容的提问来源于stack exchange,提问作者user2486276

