Python:将多列二进制数据转换为单列分类列
Alright, let's tackle this CSV compression problem step by step. You've got a 170-column dataset where 5 are unique identifiers, and the rest are binary features split across 10 categories—your goal is to shrink this down to 15 columns total (5 identifiers + 10 category columns). Here's how to do it cleanly with pandas:
核心思路
The key here is to group all binary features by their 10 categories, then aggregate each group into a single column. Since you're working with binary data, there are a few common aggregation approaches depending on what you need from the compressed data.
Step 1: Load & Prep Your Data
First, load your CSV and separate the identifier columns from the binary feature columns:
import pandas as pd # Load your actual CSV file df = pd.read_csv("your_dataset.csv") # Define your 5 unique identifier columns (match your actual column names exactly!) id_columns = ["Platform", "ID", "date", "通话时长", "name"] # Extract all binary feature columns (everything not in the identifier list) feature_columns = [col for col in df.columns if col not in id_columns]
Step 2: Group Features by Category
Next, you need to map each binary feature to its corresponding category. This depends on how your feature columns are named—most datasets use consistent prefixes (e.g., Billing_issue1, Billing_issue2 for a "Billing" category). Here's how to automate this if your columns follow a pattern:
# Example: If your 10 categories are named like "Category1", "Category2", ..., "Category10" # And feature columns start with the category name (e.g., "Category1_featureA") category_groups = {} for category in ["Category1", "Category2", "Category3", "Category4", "Category5", "Category6", "Category7", "Category8", "Category9", "Category10"]: # Grab all features that start with this category's name category_groups[category] = [col for col in feature_columns if col.startswith(category)] # If your columns don't have a consistent prefix, you'll need to manually define the groups: # category_groups = { # "Billing": ["late_fee", "payment_failed", "invoice_error"], # "Technical": ["app_crash", "login_failure", "slow_load"], # # ... add the remaining 8 categories # }
Step 3: Aggregate Each Category into One Column
Now, compress each category's features into a single column. Choose the aggregation method that fits your use case:
Option 1: Count of Active Features (Most Common)
Count how many binary features are set to 1 for each category:
# Start with just the identifier columns compressed_df = df[id_columns].copy() # Add aggregated columns for each category for category_name, feature_list in category_groups.items(): compressed_df[category_name] = df[feature_list].sum(axis=1)
Option 2: Binary Flag (Any Active Feature)
If you only care whether at least one feature in the category is 1 (instead of the count), use this:
for category_name, feature_list in category_groups.items(): compressed_df[category_name] = df[feature_list].any(axis=1).astype(int)
Option 3: List of Active Features
If you need to track which features were active (not just count/flag), concatenate the feature names:
for category_name, feature_list in category_groups.items(): # For each row, join the names of features that are 1 compressed_df[category_name] = df[feature_list].apply( lambda row: ", ".join([col for col, val in row.items() if val == 1]), axis=1 )
Step 4: Verify the Result
Check that your compressed dataframe has exactly 15 columns:
print(f"Compressed columns count: {len(compressed_df.columns)}") # Should output 15
Quick Notes
- Double-check that your
category_groupscorrectly maps all 165 feature columns to the 10 categories (no missing or misassigned columns!). - If your binary data uses values other than
0/1(e.g.,True/False), adjust the aggregation logic accordingly (e.g.,sum()still works for booleans in pandas).
内容的提问来源于stack exchange,提问作者REFER

