You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于零售交易Pandas数据集的市场篮子分析数据集构建问询

Got it, let's tackle this market basket analysis (MBA) dataset setup step by step. The core rule here is that transactions sharing the same Assistant_name and Date belong to the same "basket" (invoice)—so our goal is to group those transactions together and reshape the data into a format that's ready for MBA algorithms like Apriori.

1. First: Load and Prep the Raw Data

First off, let's load your dataset and make sure our key fields are in the right shape. For example, if your Date column includes timestamps, we'll want to strip that down to just the date part to avoid splitting baskets that happened on the same day but different times.

import pandas as pd

# Load your retail transaction data (replace with your actual file path)
df = pd.read_csv('retail_transactions.csv')

# Clean up the Date column if needed (convert to date-only if it has timestamps)
if df['Date'].dtype == 'object':
    df['Date'] = pd.to_datetime(df['Date']).dt.date  # Keep only YYYY-MM-DD
2. Create Unique Invoice IDs

Next, we need to assign a unique identifier to each group of transactions that share the same assistant and date. You have two solid options here:

  • Option 1: Numeric Invoice IDs (great for efficiency with MBA algorithms):
    We'll use groupby().ngroup() to generate a sequential number for each unique (Assistant_name, Date) pair.

    df['InvoiceID'] = df.groupby(['Assistant_name', 'Date']).ngroup()
    
  • Option 2: Human-Readable String IDs (useful for debugging or reporting):
    Combine the assistant name and date into a single string (e.g., JohnDoe_2024-05-20):

    df['InvoiceID'] = df['Assistant_name'] + '_' + df['Date'].astype(str)
    
3. Reshape into MBA-Ready Basket Format

MBA works best with data where each row represents one basket (invoice), and the items in that basket are listed together. We can do this by grouping on our new InvoiceID and aggregating the products into a list:

# Create a basket dataframe: one row per invoice, with a list of products
basket_df = df.groupby('InvoiceID')['Product'].apply(list).reset_index(name='Basket')

If you need a one-hot encoded matrix (required for some MBA libraries like mlxtend's Apriori implementation), add this step:

from mlxtend.preprocessing import TransactionEncoder

# Convert basket lists to one-hot encoded columns
te = TransactionEncoder()
one_hot_matrix = te.fit_transform(basket_df['Basket'])
one_hot_basket = pd.DataFrame(one_hot_matrix, columns=te.columns_)
Quick Check to Verify

Before diving into analysis, double-check that your grouping worked correctly. Run this to spot-test a few invoices:

# Pick a random InvoiceID and see all its transactions
sample_invoice = basket_df['InvoiceID'].sample(1).values[0]
print(f"Transactions for Invoice {sample_invoice}:")
print(df[df['InvoiceID'] == sample_invoice][['Assistant_name', 'Date', 'Product']])

That should give you a clean, ready-to-use dataset for market basket analysis—whether you're looking for frequent item sets, association rules, or customer behavior patterns.

内容的提问来源于stack exchange,提问作者Alessandro Ceccarelli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:51:43