You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

SKLearn:如何获取OneHotEncoder转换后的数据集特征名称?

Getting Transformed Feature Names from scikit-learn OneHotEncoder

Let's break down how to map your original features to the one-hot encoded feature names, using your example data and leveraging the active_features_, feature_indices_, and n_values_ properties of OneHotEncoder.

First, Let's Set Up the Data and Encoder

First, let's replicate your example (note: as_matrix() is deprecated in pandas, so we'll use .values instead):

import pandas as pd
from sklearn.preprocessing import OneHotEncoder

# Your input data
df = pd.DataFrame({"a": [0, 1, 2, 0], "b": [0, 1, 4, 5], "c": [0, 1, 4, 5]})
data = df.values

# Initialize the encoder (sparse=False lets us see the encoded data easily)
encoder = OneHotEncoder(sparse=False, handle_unknown='ignore')
encoded_data = encoder.fit_transform(data)

Understanding the Key Properties

Before we generate the names, let's clarify what those properties mean with your data:

  • encoder.n_values_: Array showing the number of "possible categories" for each original feature. By default, this is calculated as max(feature_values) + 1. For your data, this returns array([3, 6, 6]) (since a max is 2, b/c max is 5).
  • encoder.feature_indices_: Array marking the start index of each original feature's encoded columns. For your data, this is array([ 0, 3, 9, 15]) — meaning feature a uses indices 0-2, b uses 3-8, c uses 9-14.
  • encoder.active_features_: Array of global indices for categories that actually appear in your data. For your example, this returns array([ 0, 1, 2, 3, 4, 7, 8, 9, 10, 13, 14]) — skipping indices for missing categories like b_2, b_3, c_2, c_3.

Generating the Encoded Feature Names

We can map each index in active_features_ back to its original feature and category with this code:

# Get original feature names from the DataFrame
original_features = df.columns.tolist()

encoded_feature_names = []
for active_idx in encoder.active_features_:
    # Find which original feature this encoded column belongs to
    feature_idx = next(i for i, idx in enumerate(encoder.feature_indices_) if idx > active_idx) - 1
    # Calculate the original category value (local index within the feature)
    category_value = active_idx - encoder.feature_indices_[feature_idx]
    # Build the feature name
    encoded_feature_names.append(f"{original_features[feature_idx]}_{category_value}")

# Print the result
print(encoded_feature_names)

This will output exactly the names we expect:

['a_0', 'a_1', 'a_2', 'b_0', 'b_1', 'b_4', 'b_5', 'c_0', 'c_1', 'c_4', 'c_5']

A Simpler Method for Newer scikit-learn Versions

If you're using scikit-learn 0.22 or later, the OneHotEncoder has a built-in method get_feature_names_out() that does this automatically:

# Directly get encoded feature names
encoded_feature_names = encoder.get_feature_names_out(original_features)
print(encoded_feature_names)

This returns the same list of names, without needing to manually parse the properties.

内容的提问来源于stack exchange,提问作者user7468395

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:34:37