基于Apriori算法与FPA的序列规则挖掘:Excel数据加载及挖掘方法问询
Got it, let’s break this down step by step—loading your Excel data into a pandas DataFrame and prepping it for Apriori and FP-Growth (FPA) pattern mining is straightforward once you handle the item separation correctly. Here’s how to do it:
Start by loading the file with read_excel, just like you were doing. Let’s add a quick check to make sure the data loaded as expected:
import pandas as pd # Replace with your actual file path df = pd.read_excel("your_sequence_data.xlsx") # Print the first 5 rows to confirm the column and data structure print(df.head())
You’ll see a single column (something like Sequence or whatever you named it) with entries formatted like ItemA--->ItemB--->ItemC.
Next, we need to turn each string in that column into a list of individual items—this is the format most pattern mining libraries expect. Use str.split() to break apart the ---> separators, and add a quick cleanup to remove any accidental whitespace around items:
# Replace 'Column_Name' with the actual name of your single column from Excel df['transaction_list'] = df['Column_Name'].str.split('--->').apply(lambda x: [item.strip() for item in x]) # Optional: Drop the original column if you don't need it anymore df = df.drop('Column_Name', axis=1) # Verify the split worked print(df['transaction_list'].head())
Now each row has a list of items, e.g., ['ItemA', 'ItemB', 'ItemC'].
We’ll use the mlxtend library—it has solid implementations of both Apriori and FP-Growth. First install it if you haven’t:
pip install mlxtend
Then convert your transaction lists into a one-hot encoded DataFrame (this is what the mining algorithms need):
from mlxtend.preprocessing import TransactionEncoder from mlxtend.frequent_patterns import apriori, fpgrowth # Convert transaction lists to one-hot encoded format te = TransactionEncoder() encoded_data = te.fit(df['transaction_list']).transform(df['transaction_list']) one_hot_df = pd.DataFrame(encoded_data, columns=te.columns_) # Check the encoded data print(one_hot_df.head())
This will give you a DataFrame where each column is an item, and rows are 1/0 flags indicating if the item is present in the transaction.
Now find frequent itemsets and generate association rules. Adjust min_support and min_threshold based on your dataset size (smaller datasets need higher support, larger ones can use lower values):
# Find frequent itemsets with Apriori frequent_itemsets_apriori = apriori(one_hot_df, min_support=0.1, use_colnames=True) # Generate association rules (adjust min_threshold as needed) from mlxtend.frequent_patterns import association_rules apriori_rules = association_rules(frequent_itemsets_apriori, metric="confidence", min_threshold=0.7) # Print key columns to inspect results print(apriori_rules[['antecedents', 'consequents', 'support', 'confidence']])
FP-Growth is usually faster than Apriori, especially for large datasets. The workflow is almost identical:
# Find frequent itemsets with FP-Growth frequent_itemsets_fpgrowth = fpgrowth(one_hot_df, min_support=0.1, use_colnames=True) # Generate association rules fpgrowth_rules = association_rules(frequent_itemsets_fpgrowth, metric="confidence", min_threshold=0.7) # Inspect results print(fpgrowth_rules[['antecedents', 'consequents', 'support', 'confidence']])
Wait, you mentioned "sequence rules"—if you’re looking for patterns where order matters (e.g., ItemA followed by ItemB), Apriori and FP-Growth won’t cut it. Those algorithms are for unordered association rules. For sequential pattern mining, you’d need tools like PrefixSpan (available in libraries like pysdm or mlxtend has some sequential support). But if you just mean finding frequent item sets regardless of order, the steps above work perfectly.
内容的提问来源于stack exchange,提问作者Shivam

