You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python Pandas其他列条件新增交易分类列的实现需求

Add transaction_category Column to Your Pandas DataFrame

Got it, let's tackle this problem step by step. The goal is to categorize each transaction based on whether it contains only category X, only category Y, or a mix of both. Here's a straightforward solution:

Full Working Code

import pandas as pd

# Initialize your DataFrame as provided
df = pd.DataFrame({
    'transaction_id': ['A123','A123','B345','B345','C567','C567','D678','D678'], 
    'product_id': [255472, 251235, 253764,257344,221577,209809,223551,290678], 
    'product_category': ['X','X','Y','Y','X','Y','Y','X']
})

# Define a helper function to categorize each transaction
def categorize_transaction(categories):
    unique_cats = set(categories)
    if unique_cats == {'X'}:
        return 'Only X'
    elif unique_cats == {'Y'}:
        return 'Only Y'
    else:
        return 'Mixed X&Y'

# Apply the function to each transaction group and merge back to the original DataFrame
df['transaction_category'] = df.groupby('transaction_id')['product_category'].transform(categorize_transaction)

# Print the result to verify
print(df)

How It Works

  1. Helper Function: The categorize_transaction function takes all product categories for a single transaction, converts them to a set (to get unique values), and returns the appropriate category label.
  2. Group & Transform: Using groupby('transaction_id') groups rows by each unique transaction ID. The transform method ensures we keep the original row count—instead of returning a condensed group result, it maps the category label back to every row in the original transaction.
  3. Assign New Column: The result of the transform is directly assigned to the new transaction_category column.

Result

After running the code, your DataFrame will look like this:

transaction_idproduct_idproduct_categorytransaction_category
A123255472XOnly X
A123251235XOnly X
B345253764YOnly Y
B345257344YOnly Y
C567221577XMixed X&Y
C567209809YMixed X&Y
D678223551YMixed X&Y
D678290678XMixed X&Y

Alternative Concise Version

If you prefer a one-liner without a separate helper function, you can use a lambda directly in the transform:

df['transaction_category'] = df.groupby('transaction_id')['product_category'].transform(
    lambda x: 'Only X' if set(x) == {'X'} else ('Only Y' if set(x) == {'Y'} else 'Mixed X&Y')
)

内容的提问来源于stack exchange,提问作者jeangelj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 04:13:01