Apriori算法问题咨询:小数据结果冗余且运行超时
Let’s break down your two issues and fix them step by step:
1. Why you’re getting way too many results
Your min_support=0.03 is far too low for a dataset with only 10 transactions. Support is calculated as (number of transactions containing the itemset) / (total transactions). For 10 transactions, 0.03 means an itemset only needs to appear once to be considered frequent. With 20 items per transaction, this leads to an explosion of possible itemset combinations—almost every possible pair (or longer set) will qualify, hence the flood of results.
2. Why the runtime is longer than expected
Again, the root cause is the overly low min_support. Apriori works by generating candidate itemsets and pruning those that don’t meet the support threshold. When support is this low, the algorithm has to generate and check a massive number of candidate sets (the count grows exponentially with item quantity), which slows things down even for small datasets.
How to fix this
Here are concrete changes to your code and workflow:
Adjust the support threshold to match your dataset size
Calculate a meaningful support value based on your transaction count. For example, if you want itemsets that appear in at least 3 out of 10 transactions, set min_support=0.3 (3/10). This will immediately prune most infrequent itemsets.
Clean up your transaction data
Your code converts every value to a string—including empty values, which become "nan" in your transactions. These noise items add unnecessary candidates. Filter them out:
import pandas as pd from apyori import apriori from collections import Counter # Load data and clean transactions dataset = pd.read_csv("data.csv", header=None) transactions = [] for i in range(0, 10): temp = [] for j in range(0, 20): item = str(dataset.values[i,j]) # Skip empty/NaN items if item != 'nan' and item.strip() != '': temp.append(item) transactions.append(temp) # Optional: Check item frequencies to pick the right support threshold all_items = [item for trans in transactions for item in trans] item_counts = Counter(all_items) print("Item frequencies:") for item, count in item_counts.items(): print(f"{item}: {count}") # Run Apriori with adjusted parameters associations = apriori( transactions, min_support=0.3, # Adjust based on your frequency check min_confidence=0.7, min_lift=3, min_length=2 ) result = list(associations) # Print trimmed results print(f"Number of associations: {len(result)}") for res in result: print(res)
Tweak other parameters if needed
min_length: If you don’t need longer itemsets (e.g., 3+ items), keep it at 2 or set it higher to reduce results.min_confidence/min_lift: These filter rules after frequent itemsets are generated, but fixing support first is the most impactful change.
内容的提问来源于stack exchange,提问作者CDR

