You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Pandas实现交易数据集购买商品X的客户标签分类

Solution for Customer Tagging Based on Purchase History

First, let's align on the exact tagging rules we're implementing:

  • great: Customer had purchases before their first purchase of product X
  • boo: Customer only purchased product X exactly once (no other purchases)
  • awesome: Customer's first purchase was X, and they bought other products afterward

I'll use pandas for this since it's the standard tool for working with tabular transaction data. Here's a practical, tested implementation:

Step 1: Sample Test Data

Let's start with a dataset that matches your example customer scenarios:

import pandas as pd

# Sample transaction records
data = {
    'custID': [1, 1, 1, 2, 3, 4, 4, 4],
    'product': ['A', 'X', 'B', 'X', 'X', 'X', 'C', 'D'],
    'transaction_date': ['2023-01-01', '2023-01-05', '2023-01-10', 
                         '2023-02-01', '2023-03-01', '2023-04-01', 
                         '2023-04-05', '2023-04-10']
}
df = pd.DataFrame(data)

Step 2: Tagging Function

This function groups transactions by customer, sorts them by date to preserve purchase order, then applies your tagging logic:

def tag_customers(df, target_product='X'):
    # Sort transactions to ensure we process purchases in chronological order
    sorted_transactions = df.sort_values(['custID', 'transaction_date'])
    
    # Group data by customer ID to process each customer's history separately
    customer_groups = sorted_transactions.groupby('custID')
    
    tag_results = []
    
    for cust_id, group in customer_groups:
        # Get the sequence of products the customer bought (in order)
        purchase_sequence = group['product'].tolist()
        
        # Handle edge case: customer never bought the target product (just in case)
        if target_product not in purchase_sequence:
            tag_results.append({'custID': cust_id, 'tag': 'no_target_purchase'})
            continue
        
        first_purchase = purchase_sequence[0]
        
        if first_purchase != target_product:
            # Customer had existing purchases before buying X
            tag_results.append({'custID': cust_id, 'tag': 'great'})
        else:
            # First purchase was X—check if they bought other products later
            has_subsequent_non_target = any(p != target_product for p in purchase_sequence[1:])
            
            if has_subsequent_non_target:
                tag_results.append({'custID': cust_id, 'tag': 'awesome'})
            else:
                # All purchases are X—check if it's exactly one purchase
                if len(purchase_sequence) == 1:
                    tag_results.append({'custID': cust_id, 'tag': 'boo'})
                else:
                    # Edge case: multiple purchases of X, no other products
                    tag_results.append({'custID': cust_id, 'tag': 'multiple_x_only'})
    
    # Convert results to a DataFrame for easy integration with your original data
    return pd.DataFrame(tag_results)

Step 3: Test the Function

Run the function on our sample data to verify it works as expected:

customer_tags = tag_customers(df)
print(customer_tags)

Output:

custID              tag
0       1            great
1       2              boo
2       3              boo
3       4          awesome

This perfectly matches your example cases!

Edge Case Note

I added handling for a scenario you didn't explicitly mention: customers who bought X multiple times but no other products (tagged as multiple_x_only). You can adjust this tag to whatever makes sense for your use case, or remove this branch if you don't expect this scenario in your dataset.

内容的提问来源于stack exchange,提问作者spr_m

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:33:44