多列共享同一类别时的Label encoding实现及Dataframe输出调整问询
Got it, let's solve this problem where you need to apply Label Encoding across multiple DataFrame columns that share the same set of categorical values. The core goal here is to make sure the exact same category gets the exact same encoded number across all target columns—avoiding the common pitfall of fitting a separate encoder for each column, which would lead to inconsistent mappings.
Step 1: Set Up Example Data
First, let's use a sample DataFrame to mimic your current setup:
import pandas as pd # Your current DataFrame (example) df = pd.DataFrame({ 'category_col1': ['apple', 'banana', 'cherry', 'apple'], 'category_col2': ['banana', 'cherry', 'apple', 'banana'], 'unrelated_col': ['x', 'y', 'z', 'y'] # Column that doesn't share the category set })
Step 2: Import Required Tools
We'll use LabelEncoder from scikit-learn for the encoding:
from sklearn.preprocessing import LabelEncoder
Step 3: Define Target Columns & Create a Unified Encoder
First, specify which columns share the category set. Then, we'll collect all unique values from these columns to train a single encoder—this ensures consistent mappings:
# List of columns that share the same category values target_columns = ['category_col1', 'category_col2'] # Combine all values from target columns to get the full set of unique categories all_unique_categories = pd.concat([df[col] for col in target_columns]).unique() # Initialize and fit the encoder on the combined category set label_encoder = LabelEncoder() label_encoder.fit(all_unique_categories)
Step 4: Apply the Encoder to Target Columns
Now, apply the pre-fitted encoder to each of your target columns:
# Transform each target column with the unified encoder for col in target_columns: df[col] = label_encoder.transform(df[col])
Resulting DataFrame
After running the code, your DataFrame will look like this (with consistent encoding across the target columns):
category_col1 category_col2 unrelated_col 0 0 1 x 1 1 2 y 2 2 0 z 3 0 1 y
Bonus: Handle Unknown Categories (For Robustness)
If you expect new, unseen categories in future data (e.g., during inference), use OrdinalEncoder instead—it lets you define how to handle unknown values:
from sklearn.preprocessing import OrdinalEncoder # Initialize encoder with unknown value handling ordinal_encoder = OrdinalEncoder( categories=[all_unique_categories], handle_unknown='use_encoded_value', unknown_value=-1 # Assign -1 to any unseen category ) # Apply to target columns for col in target_columns: df[col] = ordinal_encoder.fit_transform(df[[col]])
This way, any category not in your original set will be encoded as -1 instead of throwing an error.
内容的提问来源于stack exchange,提问作者Martin

