电商不平衡二分类任务(买家为少数类)需关注的指标咨询
Great question—imbalanced classification in e-commerce (especially when the minority class is your core focus) is super common, but easy to misjudge if you just stick to basic accuracy. Let’s break down the metrics you should prioritize, tailored exactly to your scenario where identifying Class 1 (buyers) is the top goal:
Recall (True Positive Rate)
This is your non-negotiable top metric. Recall measures how many actual buyers (Class 1) your model correctly identifies, calculated asRecall = TP / (TP + FN). Since FN (misclassifying a buyer as non-buyer) directly means missing out on potential conversions or valuable customer engagement, maximizing recall ensures you don’t overlook the core group you care about. If your model has low recall, you’re leaving money on the table by ignoring real buyers.Precision
Precision tells you how many of your "predicted buyers" are actually real buyers:Precision = TP / (TP + FP). You mentioned FP (misclassifying non-buyers as buyers) hurts precision—and this matters because it leads to wasted resources (like sending discount codes to people who won’t buy) or annoying users with irrelevant outreach. If your business has high costs associated with false positives (e.g., expensive personalized campaigns), precision becomes just as critical as recall.F1-Score
When you need a balanced view of recall and precision, the F1-score is your go-to. It’s the harmonic mean of the two metrics:F1 = 2 * (Precision * Recall) / (Precision + Recall). This is perfect if you can’t afford to prioritize one over the other—say, you want to minimize both missed buyers and wasted marketing spend.AUC-PR (Area Under the Precision-Recall Curve)
While AUC-ROC is a popular overall metric, it can be misleading for imbalanced datasets (since it weights majority-class performance heavily). AUC-PR focuses exclusively on the positive class (buyers), measuring how well your model maintains precision as recall increases. A high AUC-PR means your model does a good job identifying buyers without flooding you with false positives.Confusion Matrix
Not a single metric, but an absolute must for context. A confusion matrix lays out exact counts of TP, TN, FP, and FN, so you can see exactly where your model is failing. For example: Are you missing more buyers (high FN) or mislabeling too many non-buyers (high FP)? This guides your next steps—like adjusting classification thresholds or tweaking training data.Cost-Weighted Metrics
If you can assign concrete business costs to each error type (e.g., a FN costs $50 in lost revenue, a FP costs $10 in wasted campaign spend), use cost-sensitive metrics. You can calculate an expected cost for your model, or use weighted versions of recall/precision that penalize the more costly error more heavily. This aligns your model’s performance directly with your business goals.
Critical Note: Avoid Accuracy at All Costs
Accuracy ((TP + TN) / Total Samples) is useless here. If 95% of your data is non-buyers, a model that predicts "non-buyer" for every single sample will have 95% accuracy—but it will completely fail at identifying your core Class 1 buyers. Always ignore accuracy for imbalanced classification tasks.
内容的提问来源于stack exchange,提问作者thanasissdr

