如何对市场篮项目进行聚类划分?含数据集及需求说明
Hey there! Let's work through how to cluster your items based on their co-occurrence frequency in transactions—this is a classic problem in market basket analysis, so we've got a straightforward workflow to follow.
First, we need to build a matrix where each entry (i,j) represents how many times item Vi and Vj appear together in the same transaction (i.e., both have a value of 1).
If your data is in a tabular format (like a CSV with transactions as rows and items as columns), we can use pandas to calculate this easily. The transpose-dot product trick is perfect here: when you transpose the item matrix and multiply it by itself, each cell ends up being the count of co-occurrences between two items.
Here's a quick snippet to do this:
import pandas as pd import numpy as np # Load your transaction data (replace with your actual file path) df = pd.read_csv("transactions.csv", header=0) # Calculate co-occurrence matrix: df.T.dot(df) gives co-occurrences co_occur_matrix = df.T.dot(df) # Optional: Set diagonal to 0 if you don't want to count an item co-occurring with itself np.fill_diagonal(co_occur_matrix.values, 0)
The resulting co_occur_matrix for your sample data would look like this:
| V1 | V2 | V3 | V4 | |
|---|---|---|---|---|
| V1 | 0 | 2 | 3 | 2 |
| V2 | 2 | 0 | 4 | 3 |
| V3 | 3 | 4 | 0 | 2 |
| V4 | 2 | 3 | 2 | 0 |
Since we're working with similarity (co-occurrence counts), we need a clustering method that can use this data directly, or convert it to a distance metric. Two great options are:
- Hierarchical Agglomerative Clustering: Perfect if you want to visualize how items group together (via a dendrogram) and decide on k later. It starts with each item as its own cluster and merges the most similar pairs iteratively.
- K-Means Clustering: Ideal if you already know the number of clusters k. Note that K-Means works with distance metrics, so we'll need to convert our co-occurrence matrix to a distance matrix first (e.g., higher co-occurrence = lower distance).
Let's cover both approaches with code.
Option A: Hierarchical Clustering
from scipy.cluster.hierarchy import linkage, dendrogram import matplotlib.pyplot as plt # Convert co-occurrence matrix to a condensed distance matrix (linkage expects this format) # Use 1/(1 + co_occurrence) to turn high co-occurrence into low distance distance_matrix = 1 / (1 + co_occur_matrix.values) condensed_distance = distance_matrix[np.triu_indices_from(distance_matrix, k=1)] # Perform hierarchical clustering linked = linkage(condensed_distance, method='ward') # 'ward' minimizes variance within clusters # Plot dendrogram to visualize cluster relationships plt.figure(figsize=(10, 6)) dendrogram(linked, labels=co_occur_matrix.index, orientation='top', distance_sort='descending', show_leaf_counts=True) plt.title('Item Clustering Dendrogram') plt.xlabel('Items') plt.ylabel('Distance') plt.show()
From the dendrogram, you can cut the tree at a certain height to get k clusters that make sense for your data.
Option B: K-Means Clustering
from sklearn.cluster import KMeans from sklearn.preprocessing import normalize # Convert co-occurrence matrix to a distance matrix: max co-occurrence minus current value max_co_occur = co_occur_matrix.values.max() distance_matrix = max_co_occur - co_occur_matrix.values # Normalize the distance matrix to ensure scale doesn't skew clustering normalized_dist = normalize(distance_matrix) # Initialize K-Means with your chosen k k = 2 # Replace with your desired number of clusters kmeans = KMeans(n_clusters=k, random_state=42) clusters = kmeans.fit_predict(normalized_dist) # Map items to their clusters and print results item_clusters = pd.DataFrame({'Item': co_occur_matrix.index, 'Cluster': clusters}) print(item_clusters)
If you don't know k, use the elbow method: plot the inertia (sum of squared distances) for different k values and pick the k where the inertia starts to decrease slowly.
- Handling Large Datasets: If your dataset is huge (thousands of items/transactions), use sparse matrices (via
scipy.sparse) to save memory when calculating the co-occurrence matrix. - Normalized Similarity: Instead of raw co-occurrence counts, you could use Jaccard similarity (proportion of transactions where both items appear, divided by transactions where either appears) for a more normalized measure of co-occurrence strength.
内容的提问来源于stack exchange,提问作者BS.Mira

