Pandas多特征对应多标签处理:基于众数生成唯一标签映射
Hey there! No worries at all—every Pandas user starts out asking these kinds of questions, so yours is totally valid and totally worth answering 😊
Here's a straightforward way to create that unique label mapping based on the mode of label for each feature2 + feature3 combination:
Step 1: Set up your data (if you haven't already)
First, let's get your sample data into a Pandas DataFrame:
import pandas as pd # Your sample data data = { 'feature1': ['a', 'b', 'c'], 'feature2': [2, 2, 2], 'feature3': [3, 3, 3], 'label': [1, 1, 0] } df = pd.DataFrame(data)
Step 2: Calculate mode per feature combination
We'll group the data by feature2 and feature3, then compute the mode of the label column for each group. This gives us the most frequent label for each combination:
# Calculate mode for each (feature2, feature3) pair # Using lambda to handle cases where multiple modes exist (we pick the first one) label_mapping = df.groupby(['feature2', 'feature3'])['label'].agg(lambda x: x.mode().iloc[0]).reset_index()
Step 3: Convert to a lookup dictionary (optional but handy)
If you want a quick way to look up the label for any feature2 + feature3 pair, convert the mapping to a dictionary where the key is a tuple of the two features:
mapping_dict = label_mapping.set_index(['feature2', 'feature3'])['label'].to_dict()
For your sample data, this dictionary will look like:
{(2, 3): 1}
Step 4: Apply the mapping to your data (if needed)
If you want to add this mapped label back to your original DataFrame, you can do this with apply:
df['mapped_label'] = df.apply(lambda row: mapping_dict[(row['feature2'], row['feature3'])], axis=1)
Your updated DataFrame will now have a mapped_label column with the mode-based label for each row:
| feature1 | feature2 | feature3 | label | mapped_label |
|---|---|---|---|---|
| a | 2 | 3 | 1 | 1 |
| b | 2 | 3 | 1 | 1 |
| c | 2 | 3 | 0 | 1 |
Notes on edge cases
- If a
feature2+feature3combination has multiple modes (e.g., two labels that appear equally often), thelambda x: x.mode().iloc[0]will pick the first one in the sorted order. If you need a different behavior (like picking the smallest label, or handling it explicitly), you can adjust the lambda function accordingly. - Since you mentioned
feature1is a unique ID with no actual use, you can drop it anytime withdf = df.drop('feature1', axis=1)if it's cluttering your data.
Hope this helps you out! Feel free to follow up if you run into any snags with more complex data or edge cases.
内容的提问来源于stack exchange,提问作者Kevin

