基于杂货采购历史数据识别特定条件下异常高频购买商品的ML方案
Hey there! Let's walk through how to solve this problem of spotting abnormally frequent purchases when customers shop online and opt for extra shipping. I’ll break down the best algorithms to use and a step-by-step implementation approach:
1. First: Clarify the Problem & Prep Your Data
Before jumping into algorithms, let’s align on what we’re targeting: we want products that are purchased significantly more often in the subset of orders where customers shopped online and paid extra shipping, compared to either the overall purchase history or other subsets (like online without extra shipping, or in-store).
First, get your data ready with these cleaned key fields:
user_id: Unique customer identifierproduct_id: Unique product identifierorder_date: Purchase datechannel: Shopping channel (online/offline)extra_shipping: Boolean (True if customer paid extra for shipping)quantity: Number of units purchased
Then filter and aggregate core metrics per product:
- Isolate the target subset:
online = True+extra_shipping = True - Calculate metrics like:
target_purchase_count: Total times the product was bought in the target subsettarget_user_count: Unique customers who bought it in the target subsetoverall_purchase_count: Total purchases across all channels/shipping optionspurchase_ratio:target_purchase_count / overall_purchase_count(captures overrepresentation in the target subset)
2. Algorithm Options (From Simple to Advanced)
A. Statistical Methods (Quick, Interpretable Results)
Perfect for straightforward, easy-to-explain insights without heavy ML:
- Z-Score/Modified Z-Score:
- Calculate how many standard deviations a product’s
target_purchase_countis from the mean of all products in the target subset. Products with a Z-score > 3 (or your chosen threshold) are outliers (abnormally frequent). - Use Modified Z-Score if your data is skewed (common with purchase frequency) since it’s robust to outliers.
- Calculate how many standard deviations a product’s
- Chi-Squared Test:
- Build a contingency table for each product comparing purchases in the target subset vs. all other orders. The test will flag products where the purchase distribution is significantly different from expected. Low p-values (<0.05) combined with a higher-than-expected share in the target subset mark abnormal products.
B. Unsupervised ML Algorithms (For Hidden Patterns)
Great for nuanced data relationships (e.g., category-specific behavior):
- Isolation Forest:
- Built specifically for anomaly detection. It isolates outliers (products with extremely high target subset frequency) by randomly splitting feature space—no labeled data needed!
- Feed it features like
target_purchase_count,purchase_ratio, andtarget_user_count; the model assigns an anomaly score, with-1marking outliers.
- Local Outlier Factor (LOF):
- Calculates each product’s density relative to its neighbors. Products with far higher frequency than similar items (e.g., same category) get a high LOF score, flagging them as anomalies.
- K-Means Clustering:
- Cluster products by their purchase metrics. Look for small clusters of products with drastically higher
target_purchase_countthan the rest—these are your abnormally frequent items.
- Cluster products by their purchase metrics. Look for small clusters of products with drastically higher
C. Supervised ML (If You Have Labeled Data)
If you already know which products are "abnormally frequent" (e.g., from past audits), train a classification model:
- Use algorithms like Random Forest or Logistic Regression, where the target label is
is_abnormal(True/False) and features are purchase metrics + product attributes (category, price, etc.). - Ideal for scaling detection across large datasets and incorporating more context.
3. Step-by-Step Implementation (Python Example)
Let’s use Pandas and Scikit-learn to put this into action:
Preprocess the Data
import pandas as pd # Load purchase history data df = pd.read_csv('grocery_purchase_history.csv') # Filter target subset: online + paid extra shipping target_subset = df[(df['channel'] == 'online') & (df['extra_shipping'] == True)] # Calculate product-level metrics product_metrics = target_subset.groupby('product_id').agg( target_purchase_count=('order_id', 'count'), target_user_count=('user_id', 'nunique') ).reset_index() # Add overall purchase count for ratio calculation overall_counts = df.groupby('product_id')['order_id'].count().reset_index(name='overall_purchase_count') product_metrics = product_metrics.merge(overall_counts, on='product_id', how='left') product_metrics['purchase_ratio'] = product_metrics['target_purchase_count'] / product_metrics['overall_purchase_count']
Option 1: Z-Score Method
import scipy.stats as stats # Calculate Z-score for target purchase count product_metrics['z_score'] = stats.zscore(product_metrics['target_purchase_count']) # Filter abnormal items (Z-score > 3) abnormal_products_zscore = product_metrics[product_metrics['z_score'] > 3]
Option 2: Isolation Forest
from sklearn.ensemble import IsolationForest # Select features for the model features = product_metrics[['target_purchase_count', 'purchase_ratio', 'target_user_count']] # Initialize model (contamination = expected share of anomalies, adjust based on your data) iso_forest = IsolationForest(contamination=0.05, random_state=42) product_metrics['is_anomaly'] = iso_forest.fit_predict(features) # Filter anomalies (-1 = anomaly) abnormal_products_iso = product_metrics[product_metrics['is_anomaly'] == -1]
4. Pro Tips for Better Results
- Normalize by Category: Calculate metrics per product category (e.g., compare milk to other dairy items, not all groceries) to avoid flagging universally high-demand products as abnormal.
- Validate with Business Context: Cross-check results with domain knowledge—if a product is a known impulse buy customers often pay extra shipping for, it might be expected, not an anomaly.
- Track Over Time: Run this analysis periodically to spot trends (e.g., a product suddenly becomes abnormally frequent after a promotion).
内容的提问来源于stack exchange,提问作者John legend2

