如何基于商品-客户邮箱字典生成匹配数DataFrame用于关联分析
Got it, let's build that symmetric DataFrame you need. The key idea is to calculate the number of overlapping customers between every pair of products (including each product with itself). Here's how to do it step by step:
First, let's start with your existing customer_emails dictionary. We'll use pandas to create the DataFrame, so make sure you have it imported.
Full Solution Code
import pandas as pd # Your existing customer email dictionary customer_emails = { 'Backpack': ['customer1@gmail.com', 'customer2@gmail.com', 'customer3@yahoo.com', 'customer4@msn.com'], 'Baseball Bat': ['customer1@gmail.com', 'customer3@yahoo.com', 'customer5@gmail.com'], 'Gloves': ['customer2@gmail.com', 'customer3@yahoo.com', 'customer4@msn.com'] } # Get the list of all product names to use as index/columns products = list(customer_emails.keys()) # Initialize an empty DataFrame with products as both rows and columns product_overlap = pd.DataFrame(index=products, columns=products) # Iterate over every pair of products to calculate common customers for product_a in products: for product_b in products: # Convert email lists to sets for fast intersection common_customers = set(customer_emails[product_a]) & set(customer_emails[product_b]) # Assign the count of common customers to the DataFrame cell product_overlap.loc[product_a, product_b] = len(common_customers) # Convert values to integers (optional but cleans up the output) product_overlap = product_overlap.astype(int) print(product_overlap)
What This Does
- Extract Product Names: We take the keys from your dictionary to use as both the row and column labels for our symmetric DataFrame.
- Initialize Empty DataFrame: Creates a grid where each cell will hold the count of overlapping customers between two products.
- Calculate Overlaps: For every pair of products (including a product with itself), we convert the email lists to sets (since set intersections are efficient) and count how many emails are in both sets.
- Clean Up: Converting the values to integers ensures the output looks clean (no float types).
Output
Running this code will give you exactly the symmetric table you wanted:
Backpack Baseball Bat Gloves Backpack 4 2 3 Baseball Bat 2 3 1 Gloves 3 1 3
This approach scales well too—if you add more products to your dictionary later, the code will automatically include them in the DataFrame without any extra work.
内容的提问来源于stack exchange,提问作者Trevor Theodore
相关产品推荐
相关产品推荐

