Sklearn中accuracy_score与ROC曲线输出结果差异的原因排查
First, let's check your code for any potential issues, then dive into the core reason for the discrepancy between accuracy and AUC.
1. Quick Check on Your Code
Looking at your code, you re-fit the Logistic Regression model when calculating y_scores_lr:
y_scores_lr = LogReg.fit(X_train, y_train).decision_function(X_test)
Since you set random_state=0 in train_test_split and the default Logistic Regression solver (lbfgs) has no inherent randomness (assuming you didn't tweak solver settings), re-fitting here will produce the exact same model as your initial LogReg.fit(X_train, y_train). This step is redundant but doesn't cause the score difference you're seeing. A cleaner approach would be to use the already trained model:
y_scores_lr = LogReg.decision_function(X_test)
2. Core Reason: Accuracy vs. AUC Measure Different Things
The key issue here is that accuracy and AUC are fundamentally distinct metrics, designed to evaluate separate aspects of your model:
Accuracy Score: This measures the proportion of correctly classified samples using a fixed threshold (Logistic Regression uses 0.5 by default). The formula is:
Accuracy = (True Positives + True Negatives) / Total SamplesIt only cares about how well the model performs at one specific cutoff point for labeling samples as positive or negative.
AUC (Area Under ROC Curve): This measures the model's overall ability to distinguish between positive and negative classes across all possible classification thresholds. It represents the probability that a randomly selected positive sample will have a higher model score than a randomly selected negative sample.
ROC curves plot the True Positive Rate (TPR) against the False Positive Rate (FPR) as you adjust the classification threshold—AUC is the area under this curve, capturing the model's ranking power rather than performance at a single threshold.
3. Why Your Scores Differ
Your model's AUC (0.87) indicates it has strong overall ability to rank positive samples higher than negative ones. However, the accuracy (0.84) reflects how well it classifies samples when using the default 0.5 threshold. These two values don't need to match:
- For example, if your dataset has a slight class imbalance, or if the model's optimal threshold (for maximizing accuracy) isn't 0.5, you'd naturally see a gap between the two metrics.
- AUC focuses on ranking quality, while accuracy focuses on hard classification at one cutoff. A model can have great ranking ability (high AUC) but not be perfectly calibrated to the 0.5 threshold for maximum accuracy.
4. Quick Validation
To confirm your AUC calculation is correct, you can replace decision_function(X_test) with predict_proba(X_test)[:, 1] (the predicted probabilities for the positive class). Since decision_function outputs log-odds (a monotonic transformation of probabilities), the ROC curve and AUC will be identical—this is a simple way to verify your code is working as expected.
内容的提问来源于stack exchange,提问作者StatsUser

