You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对市场篮项目进行聚类划分?含数据集及需求说明

Hey there! Let's work through how to cluster your items based on their co-occurrence frequency in transactions—this is a classic problem in market basket analysis, so we've got a straightforward workflow to follow.

Step 1: Compute the Co-occurrence Matrix

First, we need to build a matrix where each entry (i,j) represents how many times item Vi and Vj appear together in the same transaction (i.e., both have a value of 1).

If your data is in a tabular format (like a CSV with transactions as rows and items as columns), we can use pandas to calculate this easily. The transpose-dot product trick is perfect here: when you transpose the item matrix and multiply it by itself, each cell ends up being the count of co-occurrences between two items.

Here's a quick snippet to do this:

import pandas as pd
import numpy as np

# Load your transaction data (replace with your actual file path)
df = pd.read_csv("transactions.csv", header=0)

# Calculate co-occurrence matrix: df.T.dot(df) gives co-occurrences
co_occur_matrix = df.T.dot(df)

# Optional: Set diagonal to 0 if you don't want to count an item co-occurring with itself
np.fill_diagonal(co_occur_matrix.values, 0)

The resulting co_occur_matrix for your sample data would look like this:

V1V2V3V4
V10232
V22043
V33402
V42320
Step 2: Choose a Clustering Algorithm

Since we're working with similarity (co-occurrence counts), we need a clustering method that can use this data directly, or convert it to a distance metric. Two great options are:

  • Hierarchical Agglomerative Clustering: Perfect if you want to visualize how items group together (via a dendrogram) and decide on k later. It starts with each item as its own cluster and merges the most similar pairs iteratively.
  • K-Means Clustering: Ideal if you already know the number of clusters k. Note that K-Means works with distance metrics, so we'll need to convert our co-occurrence matrix to a distance matrix first (e.g., higher co-occurrence = lower distance).
Step 3: Implement Clustering

Let's cover both approaches with code.

Option A: Hierarchical Clustering

from scipy.cluster.hierarchy import linkage, dendrogram
import matplotlib.pyplot as plt

# Convert co-occurrence matrix to a condensed distance matrix (linkage expects this format)
# Use 1/(1 + co_occurrence) to turn high co-occurrence into low distance
distance_matrix = 1 / (1 + co_occur_matrix.values)
condensed_distance = distance_matrix[np.triu_indices_from(distance_matrix, k=1)]

# Perform hierarchical clustering
linked = linkage(condensed_distance, method='ward') # 'ward' minimizes variance within clusters

# Plot dendrogram to visualize cluster relationships
plt.figure(figsize=(10, 6))
dendrogram(linked,
           labels=co_occur_matrix.index,
           orientation='top',
           distance_sort='descending',
           show_leaf_counts=True)
plt.title('Item Clustering Dendrogram')
plt.xlabel('Items')
plt.ylabel('Distance')
plt.show()

From the dendrogram, you can cut the tree at a certain height to get k clusters that make sense for your data.

Option B: K-Means Clustering

from sklearn.cluster import KMeans
from sklearn.preprocessing import normalize

# Convert co-occurrence matrix to a distance matrix: max co-occurrence minus current value
max_co_occur = co_occur_matrix.values.max()
distance_matrix = max_co_occur - co_occur_matrix.values

# Normalize the distance matrix to ensure scale doesn't skew clustering
normalized_dist = normalize(distance_matrix)

# Initialize K-Means with your chosen k
k = 2 # Replace with your desired number of clusters
kmeans = KMeans(n_clusters=k, random_state=42)
clusters = kmeans.fit_predict(normalized_dist)

# Map items to their clusters and print results
item_clusters = pd.DataFrame({'Item': co_occur_matrix.index, 'Cluster': clusters})
print(item_clusters)

If you don't know k, use the elbow method: plot the inertia (sum of squared distances) for different k values and pick the k where the inertia starts to decrease slowly.

Key Notes
  • Handling Large Datasets: If your dataset is huge (thousands of items/transactions), use sparse matrices (via scipy.sparse) to save memory when calculating the co-occurrence matrix.
  • Normalized Similarity: Instead of raw co-occurrence counts, you could use Jaccard similarity (proportion of transactions where both items appear, divided by transactions where either appears) for a more normalized measure of co-occurrence strength.

内容的提问来源于stack exchange,提问作者BS.Mira

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:30:23