SKLearn:如何获取OneHotEncoder转换后的数据集特征名称?
Let's break down how to map your original features to the one-hot encoded feature names, using your example data and leveraging the active_features_, feature_indices_, and n_values_ properties of OneHotEncoder.
First, Let's Set Up the Data and Encoder
First, let's replicate your example (note: as_matrix() is deprecated in pandas, so we'll use .values instead):
import pandas as pd from sklearn.preprocessing import OneHotEncoder # Your input data df = pd.DataFrame({"a": [0, 1, 2, 0], "b": [0, 1, 4, 5], "c": [0, 1, 4, 5]}) data = df.values # Initialize the encoder (sparse=False lets us see the encoded data easily) encoder = OneHotEncoder(sparse=False, handle_unknown='ignore') encoded_data = encoder.fit_transform(data)
Understanding the Key Properties
Before we generate the names, let's clarify what those properties mean with your data:
encoder.n_values_: Array showing the number of "possible categories" for each original feature. By default, this is calculated asmax(feature_values) + 1. For your data, this returnsarray([3, 6, 6])(sinceamax is 2,b/cmax is 5).encoder.feature_indices_: Array marking the start index of each original feature's encoded columns. For your data, this isarray([ 0, 3, 9, 15])— meaning featureauses indices 0-2,buses 3-8,cuses 9-14.encoder.active_features_: Array of global indices for categories that actually appear in your data. For your example, this returnsarray([ 0, 1, 2, 3, 4, 7, 8, 9, 10, 13, 14])— skipping indices for missing categories likeb_2,b_3,c_2,c_3.
Generating the Encoded Feature Names
We can map each index in active_features_ back to its original feature and category with this code:
# Get original feature names from the DataFrame original_features = df.columns.tolist() encoded_feature_names = [] for active_idx in encoder.active_features_: # Find which original feature this encoded column belongs to feature_idx = next(i for i, idx in enumerate(encoder.feature_indices_) if idx > active_idx) - 1 # Calculate the original category value (local index within the feature) category_value = active_idx - encoder.feature_indices_[feature_idx] # Build the feature name encoded_feature_names.append(f"{original_features[feature_idx]}_{category_value}") # Print the result print(encoded_feature_names)
This will output exactly the names we expect:
['a_0', 'a_1', 'a_2', 'b_0', 'b_1', 'b_4', 'b_5', 'c_0', 'c_1', 'c_4', 'c_5']
A Simpler Method for Newer scikit-learn Versions
If you're using scikit-learn 0.22 or later, the OneHotEncoder has a built-in method get_feature_names_out() that does this automatically:
# Directly get encoded feature names encoded_feature_names = encoder.get_feature_names_out(original_features) print(encoded_feature_names)
This returns the same list of names, without needing to manually parse the properties.
内容的提问来源于stack exchange,提问作者user7468395

