Apriori关联规则转列表耗时过长问题咨询:小数据集测试异常
Hey there, let's break down why you're seeing such slow performance even with just 20 samples, and how to fix it.
核心原因:不是「转列表」慢,是规则计算本身慢
First off, the delay you're noticing when converting associations to a list isn't from the list conversion itself. The apriori function from the apyori library returns a generator/iterator—meaning it doesn't compute all the rules upfront. When you call list(associations), you're triggering the full calculation of every possible rule that meets your thresholds, and that's where the time goes.
Why is this slow even with 20 samples? Look at your transaction construction: each row has 543 features turned into string items. That's an enormous number of unique items to process, and Apriori has to generate and check tons of candidate itemset combinations—even with small sample sizes, high item counts lead to a "candidate explosion" that kills performance.
Also, your current code might be including meaningless items (like empty strings or nan values if your dataset has missing data), which adds even more unnecessary items to process.
Fixes to Speed Things Up
1. Clean Up Your Transaction Data
First, trim down the number of items per transaction by filtering out useless values:
transactions = [] for i in range(20): row_items = [] for j in range(543): val = str(dataset.values[i,j]) # Adjust this filter to match your data's meaningless values if val.strip() != "" and val != "nan": row_items.append(val) # Only add non-empty transactions (empty ones don't contribute to rules) if row_items: transactions.append(row_items)
If you're working with numerical features, consider discretizing them (e.g., binning into ranges like "low", "medium", "high") instead of converting each unique number to a string—this reduces the total number of unique items.
2. Ditch apyori for a Faster Library
apyori is a pure-Python implementation that's not optimized for performance. Swap it out for one of these faster alternatives:
Option A: mlxtend
mlxtend has a highly optimized Apriori implementation that works well with pandas DataFrames:
from mlxtend.frequent_patterns import apriori, association_rules import pandas as pd # Take first 20 rows, convert to string, then one-hot encode (mlxtend's required format) df = dataset.head(20).astype(str) df_encoded = pd.get_dummies(df) # Generate frequent itemsets frequent_itemsets = apriori(df_encoded, min_support=0.004, use_colnames=True) # Generate association rules rules = association_rules(frequent_itemsets, metric="confidence", min_threshold=0.3) # Convert to list quickly rules_list = rules.values.tolist()
Option B: efficient-apriori
This library is built specifically for fast Apriori calculations:
from efficient_apriori import apriori # Use the cleaned transactions from step 1 itemsets, rules = apriori(transactions, min_support=0.004, min_confidence=0.3) # Convert rules to list in no time rules_list = list(rules)
3. Adjust Your Apriori Parameters
Your current min_support=0.004 for 20 samples means any itemset that appears just once is considered "frequent"—this will generate an insane number of itemsets and rules. Try bumping up the thresholds to reduce the workload:
- Raise
min_supportto at least0.1(so itemsets need to appear in 2+ samples) - Increase
min_confidenceand add amin_liftthreshold to filter out low-value rules
Final Notes
The biggest win will come from switching to a faster library like mlxtend or efficient-apriori, paired with cleaning up your transaction data to remove unnecessary items. These changes should make even the full 22000-record dataset manageable.
内容的提问来源于stack exchange,提问作者CDR

