You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多列共享同一类别时的Label encoding实现及Dataframe输出调整问询

How to Apply Label Encoding to Multiple Columns with Shared Categories

Got it, let's solve this problem where you need to apply Label Encoding across multiple DataFrame columns that share the same set of categorical values. The core goal here is to make sure the exact same category gets the exact same encoded number across all target columns—avoiding the common pitfall of fitting a separate encoder for each column, which would lead to inconsistent mappings.

Step 1: Set Up Example Data

First, let's use a sample DataFrame to mimic your current setup:

import pandas as pd

# Your current DataFrame (example)
df = pd.DataFrame({
    'category_col1': ['apple', 'banana', 'cherry', 'apple'],
    'category_col2': ['banana', 'cherry', 'apple', 'banana'],
    'unrelated_col': ['x', 'y', 'z', 'y']  # Column that doesn't share the category set
})

Step 2: Import Required Tools

We'll use LabelEncoder from scikit-learn for the encoding:

from sklearn.preprocessing import LabelEncoder

Step 3: Define Target Columns & Create a Unified Encoder

First, specify which columns share the category set. Then, we'll collect all unique values from these columns to train a single encoder—this ensures consistent mappings:

# List of columns that share the same category values
target_columns = ['category_col1', 'category_col2']

# Combine all values from target columns to get the full set of unique categories
all_unique_categories = pd.concat([df[col] for col in target_columns]).unique()

# Initialize and fit the encoder on the combined category set
label_encoder = LabelEncoder()
label_encoder.fit(all_unique_categories)

Step 4: Apply the Encoder to Target Columns

Now, apply the pre-fitted encoder to each of your target columns:

# Transform each target column with the unified encoder
for col in target_columns:
    df[col] = label_encoder.transform(df[col])

Resulting DataFrame

After running the code, your DataFrame will look like this (with consistent encoding across the target columns):

category_col1  category_col2 unrelated_col
0              0              1             x
1              1              2             y
2              2              0             z
3              0              1             y

Bonus: Handle Unknown Categories (For Robustness)

If you expect new, unseen categories in future data (e.g., during inference), use OrdinalEncoder instead—it lets you define how to handle unknown values:

from sklearn.preprocessing import OrdinalEncoder

# Initialize encoder with unknown value handling
ordinal_encoder = OrdinalEncoder(
    categories=[all_unique_categories],
    handle_unknown='use_encoded_value',
    unknown_value=-1  # Assign -1 to any unseen category
)

# Apply to target columns
for col in target_columns:
    df[col] = ordinal_encoder.fit_transform(df[[col]])

This way, any category not in your original set will be encoded as -1 instead of throwing an error.

内容的提问来源于stack exchange,提问作者Martin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:30:46