You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

提取内部元组并统计元素数量,绘制itemset元素数量直方图

How to Plot a Histogram of Itemset Sizes from Your Frequency Itemsets List

Got it, let's walk through exactly how to extract the itemset sizes and visualize their distribution with a histogram. Here's a step-by-step solution using Python:

Step 1: Extract the Length of Each Internal Tuple

First, we need to pull out the size of the inner tuple from each element in your freq_itemsets list. The inner tuple is always the first element of each entry, so we can use a list comprehension to quickly collect all these lengths:

# Extract the number of elements in each inner tuple
tuple_lengths = [len(item[0]) for item in freq_itemsets]

If you want to verify the counts first (e.g., how many 1-element itemsets, 2-element, etc.), you can use collections.Counter to get a quick summary:

from collections import Counter

# Count occurrences of each itemset size
size_counts = Counter(tuple_lengths)
print("Itemset size distribution:", size_counts)

Step 2: Plot the Histogram

We'll use matplotlib (or seaborn for a cleaner look) to create the histogram. Since itemset sizes are integers, we'll set up bins to align perfectly with each integer value so the histogram is easy to interpret.

Using Matplotlib:

import matplotlib.pyplot as plt

# Set up the plot
plt.figure(figsize=(10, 6))

# Create histogram with bins aligned to integer sizes
plt.hist(
    tuple_lengths,
    bins=range(min(tuple_lengths), max(tuple_lengths) + 2),
    edgecolor='black',
    align='left'
)

# Add labels and title
plt.xlabel('Number of Elements in Itemset')
plt.ylabel('Frequency (Number of Itemsets)')
plt.title('Distribution of Itemset Sizes')

# Set x-axis ticks to match each possible itemset size
plt.xticks(range(min(tuple_lengths), max(tuple_lengths) + 1))

# Show the plot
plt.show()

Using Seaborn (for a more polished look):

If you have seaborn installed, this will give you a cleaner visualization:

import seaborn as sns
import matplotlib.pyplot as plt

plt.figure(figsize=(10, 6))

# Plot histogram without KDE (since we're dealing with discrete counts)
sns.histplot(
    tuple_lengths,
    bins=range(min(tuple_lengths), max(tuple_lengths) + 2),
    kde=False,
    edgecolor='black'
)

# Customize labels and ticks
plt.xlabel('Number of Elements in Itemset')
plt.ylabel('Frequency (Number of Itemsets)')
plt.title('Distribution of Itemset Sizes')
plt.xticks(range(min(tuple_lengths), max(tuple_lengths) + 1))

plt.show()

Notes:

  • The bins=range(...) setup ensures each bin corresponds exactly to one integer itemset size (e.g., bin 1 covers 1-element itemsets, bin 2 covers 2-element, etc.). Adding +2 to the max ensures we include the largest size in the bins.
  • If your list has any empty inner tuples (unlikely based on your examples), you can filter them out by adding a condition to the list comprehension: [len(item[0]) for item in freq_itemsets if len(item[0]) > 0]

内容的提问来源于stack exchange,提问作者tushariyer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:36:38