类别权重不平衡时逻辑回归ROC绘制及AUC值异常问题求解
Hey there! Let's break down your two logistic regression ROC curve questions and fix that stubborn AUC issue you're dealing with.
1. Drawing ROC Curves for Logistic Regression with Class Weights
Whether you set custom class_weight or not, the core logic to plot an ROC curve stays the same—it all boils down to getting the model's predicted scores for the positive class, then calculating false positive rate (FPR) and true positive rate (TPR) from those scores.
Here's what you need to keep in mind:
- Use either predicted probabilities or decision scores: You can get positive-class probabilities with
predict_proba(X_test)[:, 1], or decision function values withdecision_function(X_test)—both work for ROC calculations. - Avoid re-fitting the model unnecessarily: Re-fitting will reset your model's training state, which can lead to inconsistent results (we'll cover this more in the next section).
2. Fixing the Stagnant AUC Value (Always 0.81)
If adjusting class_weight changes your confusion matrix but leaves AUC unchanged, there are a few key issues to check and fix:
Common Causes & Solutions
- You're re-fitting the model multiple times: Looking at your code, you first call
logmodel.fit(X_train,y_train), then later uselogmodel.fit(X_train, y_train).decision_function(X_test)—this re-trains the model from scratch, wiping out any prior training with custom class weights. Fit the model once, then reuse it for all predictions/scores. - Your custom class weights might not align with your data: Setting
{0:0.02,1:1}gives class 0 an extremely low weight, which might not be appropriate for your dataset's class distribution. Try usingclass_weight='balanced'instead—this lets scikit-learn automatically calculate weights based on class frequencies, which often gives more meaningful results. - AUC measures ranking, not classification thresholds: Remember that AUC evaluates how well the model can rank positive samples higher than negative ones. Adjusting
class_weightchanges the model's classification threshold (which affects confusion matrix/accuracy), but doesn't always alter the model's ability to rank samples. If your AUC stays the same, it might mean your model's ranking capability isn't changing with weight adjustments—which could be normal if your data's feature-label relationships are fixed.
Fixed & Improved Code
Here's a cleaned-up version of your code that addresses these issues:
from sklearn.linear_model import LogisticRegression from sklearn.metrics import (confusion_matrix, classification_report, accuracy_score, roc_curve, roc_auc_score) import matplotlib.pyplot as plt # Initialize model with custom weights or balanced weights logmodel = LogisticRegression(solver='liblinear', class_weight={0:0.02, 1:1}) # logmodel = LogisticRegression(solver='liblinear', class_weight='balanced') # Fit the model ONCE—reuse this trained model for all evaluations logmodel.fit(X_train, y_train) # Get hard predictions for confusion matrix/accuracy predictions = logmodel.predict(X_test) # Print evaluation metrics print("Confusion Matrix:") print(confusion_matrix(y_test, predictions)) print("\nClassification Report:") print(classification_report(y_test, predictions)) acc = accuracy_score(y_test, predictions) print(f"Accuracy Score: {acc:.3f}") # Get positive-class predicted probabilities for ROC model_probs = logmodel.predict_proba(X_test) y_score = model_probs[:, 1] # Extract probabilities for class 1 # Calculate ROC curve metrics and AUC fpr, tpr, _ = roc_curve(y_test, y_score) auc = roc_auc_score(y_test, y_score) # Plot the ROC curve plt.figure(figsize=(8, 6)) plt.plot(fpr, tpr, marker='.', label=f'ROC Curve (Area = {auc:.2f})') plt.plot([0, 1], [0, 1], 'r--', label='Random Guess Baseline') plt.xlabel('False Positive Rate') plt.ylabel('True Positive Rate') plt.title('ROC Curve for Weighted Logistic Regression') plt.legend(loc='lower right') plt.show()
Quick Notes
- If you prefer using the decision function instead of predicted probabilities, replace
y_score = model_probs[:, 1]withy_score = logmodel.decision_function(X_test)—the ROC/AUC calculation will work the same way. - If AUC still doesn't change after trying balanced weights, double-check your data: Are your features actually predictive? Is there extreme class imbalance that's limiting the model's ability to learn?
内容的提问来源于stack exchange,提问作者Arunee Sridee
相关产品推荐
相关产品推荐

