基于顾客重复购物模式的行为异常检测论文任务咨询
Hey there, let's work through this customer behavior anomaly detection problem you're stuck on for your thesis. Based on the dataset details you shared, here's a structured, practical approach to nail this final step:
Since every fidelity_card_ID has a unique shopping pattern portrait, we need to formalize what "normal" looks like for each user using their historical data. Focus on these key dimensions:
- Cluster frequency distribution: Calculate how often each shopping cluster (e.g., weekly grocery, dining purchase) appears for the user. For example, User X might have 85% of their purchases in "weekly grocery", 10% in "dining", and 5% in "clothing".
- Temporal patterns: Analyze the timing of their purchases—like average days between grocery runs, or consistent weekday preferences (e.g., always buys dining supplies on Fridays).
- Sequence patterns: Identify common orderings of clusters (e.g., User Y almost always buys home cleaning supplies 1 day after a weekly grocery run).
Anomalies are deviations from a user's normal profile, and you’ll need to tie this to your business context. Examples of anomalies could be:
- A sudden shift from high-frequency clusters: A user who normally buys groceries 90% of the time suddenly makes two consecutive clothing purchases.
- Breaks in temporal routines: A user who shops for groceries every 7 days goes 12 days without a grocery run, or shops 3 days in a row.
- Rare cluster combinations: A sequence of clusters that almost never appears in the user’s history (e.g., clothing + home cleaning supplies back-to-back).
Here are actionable methods you can use for your thesis, ranging from simple statistical checks to more advanced techniques:
Statistical Threshold Method (Great for Interpretability)
This is straightforward and easy to explain in a thesis. Calculate baseline stats for each user, then flag deviations beyond a set threshold:
import pandas as pd # Calculate cluster frequency per user user_cluster_ratios = df.groupby('fidelity_card_ID')['cluster'].value_counts(normalize=True).unstack(fill_value=0) # Flag purchases where the cluster's ratio is below 5% (adjust threshold as needed) df['cluster_anomaly'] = df.apply( lambda row: user_cluster_ratios.loc[row['fidelity_card_ID'], row['cluster']] < 0.05, axis=1 ) # Calculate average days between purchases per user, flag outliers df['shopping_date'] = pd.to_datetime(df['shopping_date']) df['days_since_last_purchase'] = df.groupby('fidelity_card_ID')['shopping_date'].diff().dt.days user_avg_interval = df.groupby('fidelity_card_ID')['days_since_last_purchase'].mean() user_std_interval = df.groupby('fidelity_card_ID')['days_since_last_purchase'].std() df['temporal_anomaly'] = df.apply( lambda row: abs(row['days_since_last_purchase'] - user_avg_interval.loc[row['fidelity_card_ID']]) > 2 * user_std_interval.loc[row['fidelity_card_ID']], axis=1 ) # Combine flags: a purchase is anomalous if either flag is True df['is_anomaly'] = df['cluster_anomaly'] | df['temporal_anomaly']
Sequence Pattern Detection
Use sequence mining to identify a user’s frequent shopping sequences, then flag sequences that don’t match these patterns. Tools like PrefixSpan (in libraries like mlxtend) can help here. For example:
- Extract all frequent sequences for a user (e.g., ["weekly grocery", "dining"] with support > 20%)
- Flag any new sequence that doesn’t appear in this frequent set.
Unsupervised Machine Learning
If you want a more robust model, use anomaly detection algorithms like Isolation Forest or One-Class SVM. Convert each user’s shopping history into features (e.g., last 5 clusters, average interval, cluster frequency vector) and train the model to spot outliers.
Don’t skip this step—it’s critical for a thesis:
- Collaborate with domain experts (e.g., retail operations staff) to confirm if flagged anomalies make business sense (e.g., a sudden shift might be due to a move, not fraud; that’s a behavior change, not an anomaly).
- Adjust thresholds or model parameters if you’re getting too many false positives/negatives. For example, if 5% cluster ratio is flagging too many rare-but-normal purchases, lower it to 3%.
内容的提问来源于stack exchange,提问作者Seymour

