标签编码(Label Encoding)将对象型数据转为数值型后,如何查看编码对应关系并获取特征名称?
Hey there! Let's work through your two key questions: how to check the exact category-to-code mappings from your label encoding, and how to set up feature names (like the Iris dataset) for your decision tree plots. I'll use your bank dataset examples to make this concrete.
1. Viewing Category-to-Code Mappings for Label Encoded Features
The problem with your original code is that you're reusing the same LabelEncoder instance for every feature, which means you lose the mapping for previous features. Instead, store each encoder in a dictionary so you can reference them later.
Here's how to adjust your code:
%matplotlib inline import matplotlib.pyplot as plt import pandas as pd from sklearn.preprocessing import LabelEncoder # Load your data bank = pd.read_csv('train_bank.csv') df = pd.DataFrame(bank) # Define your object-type features (adjust this list to match your actual columns) objList = ['Gen', 'Mar', 'Edu', 'Sel'] # Create a dictionary to hold each feature's LabelEncoder label_encoders = {} for feat in objList: le = LabelEncoder() # Fit and transform the feature df[feat] = le.fit_transform(df[feat].astype(str)) # Save the encoder for later label_encoders[feat] = le # Now check the mapping for any feature (e.g., Edu) print("Mapping for 'Edu' feature:") for code, category in enumerate(label_encoders['Edu'].classes_): print(f"{category} → {code}")
This will output exactly which category corresponds to each numeric code. If you ever need to convert numeric codes back to original categories, you can use label_encoders['Edu'].inverse_transform([0, 1]).
Bonus: Fixing the Encoding Order for 'Edu'
If you want Graduate to be 1 and Not Grad to be 0 (opposite of the default alphabetical order), use OrdinalEncoder instead—it lets you specify the category order explicitly:
from sklearn.preprocessing import OrdinalEncoder # Define the desired order for the Edu feature edu_order = ['Not Grad', 'Graduate'] # First item = 0, second = 1 ordinal_encoder = OrdinalEncoder(categories=[edu_order]) # Apply to the Edu column (note the double brackets to keep it as a DataFrame) df['Edu'] = ordinal_encoder.fit_transform(df[['Edu']]) # Verify the mapping print("Custom mapping for 'Edu' feature:") for code, category in enumerate(edu_order): print(f"{category} → {code}")
2. Setting Up Feature Names and Class Names for Decision Trees
Getting feature names is straightforward—they're just the column names of your feature dataset. For class names, use the unique values from your target variable (or the encoder classes if you encoded the target).
Here's how to integrate this with your decision tree code:
# First, split your data into features (X) and target (y) # Replace 'Target' with your actual target column name (e.g., 'Loan_Approved') X = df.drop('Target', axis=1) y = df['Target'] # Get feature names directly from the feature DataFrame columns feature_names = X.columns.tolist() # Get class names: if your target is not encoded, use unique values class_names = y.unique().tolist() # If you encoded your target with LabelEncoder, use this instead: # target_le = LabelEncoder() # y = target_le.fit_transform(y) # class_names = target_le.classes_.tolist() # Now plot your decision tree with these names from sklearn import tree fig, axes = plt.subplots(nrows=1, ncols=1, figsize=(4,4), dpi=500) tree.plot_tree(classifier.estimators_[0], feature_names=feature_names, class_names=[str(name) for name in class_names], # Convert to strings if needed filled=True); fig.savefig('rf_individualtree.png')
This will produce a decision tree plot with clear, human-readable feature and class names, just like the Iris example you referenced.
内容的提问来源于stack exchange,提问作者Alex

