You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

标签编码(Label Encoding)将对象型数据转为数值型后,如何查看编码对应关系并获取特征名称?

Solving Label Encoder Mapping and Feature Name Issues for Your Bank Dataset

Hey there! Let's work through your two key questions: how to check the exact category-to-code mappings from your label encoding, and how to set up feature names (like the Iris dataset) for your decision tree plots. I'll use your bank dataset examples to make this concrete.

1. Viewing Category-to-Code Mappings for Label Encoded Features

The problem with your original code is that you're reusing the same LabelEncoder instance for every feature, which means you lose the mapping for previous features. Instead, store each encoder in a dictionary so you can reference them later.

Here's how to adjust your code:

%matplotlib inline
import matplotlib.pyplot as plt
import pandas as pd
from sklearn.preprocessing import LabelEncoder

# Load your data
bank = pd.read_csv('train_bank.csv')
df = pd.DataFrame(bank)

# Define your object-type features (adjust this list to match your actual columns)
objList = ['Gen', 'Mar', 'Edu', 'Sel']

# Create a dictionary to hold each feature's LabelEncoder
label_encoders = {}

for feat in objList:
    le = LabelEncoder()
    # Fit and transform the feature
    df[feat] = le.fit_transform(df[feat].astype(str))
    # Save the encoder for later
    label_encoders[feat] = le

# Now check the mapping for any feature (e.g., Edu)
print("Mapping for 'Edu' feature:")
for code, category in enumerate(label_encoders['Edu'].classes_):
    print(f"{category} → {code}")

This will output exactly which category corresponds to each numeric code. If you ever need to convert numeric codes back to original categories, you can use label_encoders['Edu'].inverse_transform([0, 1]).

Bonus: Fixing the Encoding Order for 'Edu'

If you want Graduate to be 1 and Not Grad to be 0 (opposite of the default alphabetical order), use OrdinalEncoder instead—it lets you specify the category order explicitly:

from sklearn.preprocessing import OrdinalEncoder

# Define the desired order for the Edu feature
edu_order = ['Not Grad', 'Graduate']  # First item = 0, second = 1
ordinal_encoder = OrdinalEncoder(categories=[edu_order])

# Apply to the Edu column (note the double brackets to keep it as a DataFrame)
df['Edu'] = ordinal_encoder.fit_transform(df[['Edu']])

# Verify the mapping
print("Custom mapping for 'Edu' feature:")
for code, category in enumerate(edu_order):
    print(f"{category} → {code}")

2. Setting Up Feature Names and Class Names for Decision Trees

Getting feature names is straightforward—they're just the column names of your feature dataset. For class names, use the unique values from your target variable (or the encoder classes if you encoded the target).

Here's how to integrate this with your decision tree code:

# First, split your data into features (X) and target (y)
# Replace 'Target' with your actual target column name (e.g., 'Loan_Approved')
X = df.drop('Target', axis=1)
y = df['Target']

# Get feature names directly from the feature DataFrame columns
feature_names = X.columns.tolist()

# Get class names: if your target is not encoded, use unique values
class_names = y.unique().tolist()
# If you encoded your target with LabelEncoder, use this instead:
# target_le = LabelEncoder()
# y = target_le.fit_transform(y)
# class_names = target_le.classes_.tolist()

# Now plot your decision tree with these names
from sklearn import tree

fig, axes = plt.subplots(nrows=1, ncols=1, figsize=(4,4), dpi=500)
tree.plot_tree(classifier.estimators_[0], 
               feature_names=feature_names, 
               class_names=[str(name) for name in class_names],  # Convert to strings if needed
               filled=True);
fig.savefig('rf_individualtree.png')

This will produce a decision tree plot with clear, human-readable feature and class names, just like the Iris example you referenced.

内容的提问来源于stack exchange,提问作者Alex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 17:44:10