跨不同类别比例的强不平衡数据集比较F1分数的可行性探讨
Great question—this is a super common pitfall when working with imbalanced datasets, especially when their class ratios vary. Let’s break this down clearly so you can make meaningful comparisons.
Short answer: No, you shouldn’t rely on raw F1 scores for cross-dataset comparisons here. Here’s why:
F1 score is the harmonic mean of precision and recall, both of which are heavily influenced by the underlying class distribution of the dataset. Even if a model has identical true performance (e.g., same ability to identify minority class samples), its F1 score can vary drastically across datasets with different minority class proportions.
For example:
- Dataset A: 360 minority / 5600 total (~6.4% minority)
- Dataset B: 120 minority / 6400 total (~1.9% minority)
Suppose your model correctly identifies 50% of minority samples (recall = 50%) in both datasets, and misclassifies 10% of majority samples as minority.
- In Dataset A: That’s 180 true positives, ~524 misclassified majority samples. Precision = 180/(180+524) ≈ 25.7%, F1 ≈ 33.9%.
- In Dataset B: That’s 60 true positives, ~628 misclassified majority samples. Precision = 60/(60+628) ≈ 8.7%, F1 ≈ 15.2%.
Same core model performance, but F1 scores are wildly different—all because of the dataset’s class ratio. Raw F1 here doesn’t reflect the model’s true ability, just the dataset’s imbalance.
Here are reliable approaches to level the playing field:
- Use Balanced F1-Score: This variant (available in scikit-learn as
balanced_f1_score) calculates the F1 score for each class, then averages them weighted by the inverse of class frequencies. This gives more weight to the minority class, effectively neutralizing the impact of differing class ratios across datasets. It’s far more comparable than raw F1. - Focus on minority-class-specific metrics + relative improvements: Instead of looking at overall F1, track metrics like minority-class recall or precision, then compare them to dataset-specific baselines. For example:
- In Dataset B, a random guess would have ~1.9% minority recall. If your model hits 20% recall, that’s a 10x improvement.
- In Dataset A, a random guess gives ~6.4% recall; a 30% recall here is a 4.7x improvement.
These relative gains are far more meaningful than absolute F1 values.
- Use threshold-agnostic probability metrics: Metrics like AUC-PR (Precision-Recall AUC) or AUC-ROC measure your model’s ability to rank samples correctly, rather than relying on a fixed classification threshold. AUC-PR is particularly suited for imbalanced data because it focuses on the minority class’s performance, and it’s much less sensitive to class distribution shifts than F1.
- Normalize F1 against a baseline: For each dataset, calculate the F1 score of a simple baseline model (e.g., a majority-class classifier, or a random classifier). Then compare your model’s F1 as a ratio of the baseline F1 (e.g.,
model_f1 / baseline_f1). This relative performance metric accounts for the dataset’s inherent imbalance.
Raw F1 scores are misleading when comparing across imbalanced datasets with different class ratios—they reflect the dataset’s imbalance more than the model’s true ability. Opt for balanced metrics, relative improvements, or threshold-agnostic scores to make fair, meaningful comparisons.
内容的提问来源于stack exchange,提问作者philosonista

