基于零售交易Pandas数据集的市场篮子分析数据集构建问询
Got it, let's tackle this market basket analysis (MBA) dataset setup step by step. The core rule here is that transactions sharing the same Assistant_name and Date belong to the same "basket" (invoice)—so our goal is to group those transactions together and reshape the data into a format that's ready for MBA algorithms like Apriori.
First off, let's load your dataset and make sure our key fields are in the right shape. For example, if your Date column includes timestamps, we'll want to strip that down to just the date part to avoid splitting baskets that happened on the same day but different times.
import pandas as pd # Load your retail transaction data (replace with your actual file path) df = pd.read_csv('retail_transactions.csv') # Clean up the Date column if needed (convert to date-only if it has timestamps) if df['Date'].dtype == 'object': df['Date'] = pd.to_datetime(df['Date']).dt.date # Keep only YYYY-MM-DD
Next, we need to assign a unique identifier to each group of transactions that share the same assistant and date. You have two solid options here:
Option 1: Numeric Invoice IDs (great for efficiency with MBA algorithms):
We'll usegroupby().ngroup()to generate a sequential number for each unique (Assistant_name, Date) pair.df['InvoiceID'] = df.groupby(['Assistant_name', 'Date']).ngroup()Option 2: Human-Readable String IDs (useful for debugging or reporting):
Combine the assistant name and date into a single string (e.g.,JohnDoe_2024-05-20):df['InvoiceID'] = df['Assistant_name'] + '_' + df['Date'].astype(str)
MBA works best with data where each row represents one basket (invoice), and the items in that basket are listed together. We can do this by grouping on our new InvoiceID and aggregating the products into a list:
# Create a basket dataframe: one row per invoice, with a list of products basket_df = df.groupby('InvoiceID')['Product'].apply(list).reset_index(name='Basket')
If you need a one-hot encoded matrix (required for some MBA libraries like mlxtend's Apriori implementation), add this step:
from mlxtend.preprocessing import TransactionEncoder # Convert basket lists to one-hot encoded columns te = TransactionEncoder() one_hot_matrix = te.fit_transform(basket_df['Basket']) one_hot_basket = pd.DataFrame(one_hot_matrix, columns=te.columns_)
Before diving into analysis, double-check that your grouping worked correctly. Run this to spot-test a few invoices:
# Pick a random InvoiceID and see all its transactions sample_invoice = basket_df['InvoiceID'].sample(1).values[0] print(f"Transactions for Invoice {sample_invoice}:") print(df[df['InvoiceID'] == sample_invoice][['Assistant_name', 'Date', 'Product']])
That should give you a clean, ready-to-use dataset for market basket analysis—whether you're looking for frequent item sets, association rules, or customer behavior patterns.
内容的提问来源于stack exchange,提问作者Alessandro Ceccarelli

