基于Pandas实现交易数据集购买商品X的客户标签分类
First, let's align on the exact tagging rules we're implementing:
- great: Customer had purchases before their first purchase of product X
- boo: Customer only purchased product X exactly once (no other purchases)
- awesome: Customer's first purchase was X, and they bought other products afterward
I'll use pandas for this since it's the standard tool for working with tabular transaction data. Here's a practical, tested implementation:
Step 1: Sample Test Data
Let's start with a dataset that matches your example customer scenarios:
import pandas as pd # Sample transaction records data = { 'custID': [1, 1, 1, 2, 3, 4, 4, 4], 'product': ['A', 'X', 'B', 'X', 'X', 'X', 'C', 'D'], 'transaction_date': ['2023-01-01', '2023-01-05', '2023-01-10', '2023-02-01', '2023-03-01', '2023-04-01', '2023-04-05', '2023-04-10'] } df = pd.DataFrame(data)
Step 2: Tagging Function
This function groups transactions by customer, sorts them by date to preserve purchase order, then applies your tagging logic:
def tag_customers(df, target_product='X'): # Sort transactions to ensure we process purchases in chronological order sorted_transactions = df.sort_values(['custID', 'transaction_date']) # Group data by customer ID to process each customer's history separately customer_groups = sorted_transactions.groupby('custID') tag_results = [] for cust_id, group in customer_groups: # Get the sequence of products the customer bought (in order) purchase_sequence = group['product'].tolist() # Handle edge case: customer never bought the target product (just in case) if target_product not in purchase_sequence: tag_results.append({'custID': cust_id, 'tag': 'no_target_purchase'}) continue first_purchase = purchase_sequence[0] if first_purchase != target_product: # Customer had existing purchases before buying X tag_results.append({'custID': cust_id, 'tag': 'great'}) else: # First purchase was X—check if they bought other products later has_subsequent_non_target = any(p != target_product for p in purchase_sequence[1:]) if has_subsequent_non_target: tag_results.append({'custID': cust_id, 'tag': 'awesome'}) else: # All purchases are X—check if it's exactly one purchase if len(purchase_sequence) == 1: tag_results.append({'custID': cust_id, 'tag': 'boo'}) else: # Edge case: multiple purchases of X, no other products tag_results.append({'custID': cust_id, 'tag': 'multiple_x_only'}) # Convert results to a DataFrame for easy integration with your original data return pd.DataFrame(tag_results)
Step 3: Test the Function
Run the function on our sample data to verify it works as expected:
customer_tags = tag_customers(df) print(customer_tags)
Output:
custID tag 0 1 great 1 2 boo 2 3 boo 3 4 awesome
This perfectly matches your example cases!
Edge Case Note
I added handling for a scenario you didn't explicitly mention: customers who bought X multiple times but no other products (tagged as multiple_x_only). You can adjust this tag to whatever makes sense for your use case, or remove this branch if you don't expect this scenario in your dataset.
内容的提问来源于stack exchange,提问作者spr_m

