基于Python Pandas其他列条件新增交易分类列的实现需求
Add
transaction_category Column to Your Pandas DataFrame Got it, let's tackle this problem step by step. The goal is to categorize each transaction based on whether it contains only category X, only category Y, or a mix of both. Here's a straightforward solution:
Full Working Code
import pandas as pd # Initialize your DataFrame as provided df = pd.DataFrame({ 'transaction_id': ['A123','A123','B345','B345','C567','C567','D678','D678'], 'product_id': [255472, 251235, 253764,257344,221577,209809,223551,290678], 'product_category': ['X','X','Y','Y','X','Y','Y','X'] }) # Define a helper function to categorize each transaction def categorize_transaction(categories): unique_cats = set(categories) if unique_cats == {'X'}: return 'Only X' elif unique_cats == {'Y'}: return 'Only Y' else: return 'Mixed X&Y' # Apply the function to each transaction group and merge back to the original DataFrame df['transaction_category'] = df.groupby('transaction_id')['product_category'].transform(categorize_transaction) # Print the result to verify print(df)
How It Works
- Helper Function: The
categorize_transactionfunction takes all product categories for a single transaction, converts them to a set (to get unique values), and returns the appropriate category label. - Group & Transform: Using
groupby('transaction_id')groups rows by each unique transaction ID. Thetransformmethod ensures we keep the original row count—instead of returning a condensed group result, it maps the category label back to every row in the original transaction. - Assign New Column: The result of the transform is directly assigned to the new
transaction_categorycolumn.
Result
After running the code, your DataFrame will look like this:
| transaction_id | product_id | product_category | transaction_category |
|---|---|---|---|
| A123 | 255472 | X | Only X |
| A123 | 251235 | X | Only X |
| B345 | 253764 | Y | Only Y |
| B345 | 257344 | Y | Only Y |
| C567 | 221577 | X | Mixed X&Y |
| C567 | 209809 | Y | Mixed X&Y |
| D678 | 223551 | Y | Mixed X&Y |
| D678 | 290678 | X | Mixed X&Y |
Alternative Concise Version
If you prefer a one-liner without a separate helper function, you can use a lambda directly in the transform:
df['transaction_category'] = df.groupby('transaction_id')['product_category'].transform( lambda x: 'Only X' if set(x) == {'X'} else ('Only Y' if set(x) == {'Y'} else 'Mixed X&Y') )
内容的提问来源于stack exchange,提问作者jeangelj
相关产品推荐
相关产品推荐

