You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

类别权重不平衡时逻辑回归ROC绘制及AUC值异常问题求解

Hey there! Let's break down your two logistic regression ROC curve questions and fix that stubborn AUC issue you're dealing with.

1. Drawing ROC Curves for Logistic Regression with Class Weights

Whether you set custom class_weight or not, the core logic to plot an ROC curve stays the same—it all boils down to getting the model's predicted scores for the positive class, then calculating false positive rate (FPR) and true positive rate (TPR) from those scores.

Here's what you need to keep in mind:

  • Use either predicted probabilities or decision scores: You can get positive-class probabilities with predict_proba(X_test)[:, 1], or decision function values with decision_function(X_test)—both work for ROC calculations.
  • Avoid re-fitting the model unnecessarily: Re-fitting will reset your model's training state, which can lead to inconsistent results (we'll cover this more in the next section).
2. Fixing the Stagnant AUC Value (Always 0.81)

If adjusting class_weight changes your confusion matrix but leaves AUC unchanged, there are a few key issues to check and fix:

Common Causes & Solutions

  • You're re-fitting the model multiple times: Looking at your code, you first call logmodel.fit(X_train,y_train), then later use logmodel.fit(X_train, y_train).decision_function(X_test)—this re-trains the model from scratch, wiping out any prior training with custom class weights. Fit the model once, then reuse it for all predictions/scores.
  • Your custom class weights might not align with your data: Setting {0:0.02,1:1} gives class 0 an extremely low weight, which might not be appropriate for your dataset's class distribution. Try using class_weight='balanced' instead—this lets scikit-learn automatically calculate weights based on class frequencies, which often gives more meaningful results.
  • AUC measures ranking, not classification thresholds: Remember that AUC evaluates how well the model can rank positive samples higher than negative ones. Adjusting class_weight changes the model's classification threshold (which affects confusion matrix/accuracy), but doesn't always alter the model's ability to rank samples. If your AUC stays the same, it might mean your model's ranking capability isn't changing with weight adjustments—which could be normal if your data's feature-label relationships are fixed.

Fixed & Improved Code

Here's a cleaned-up version of your code that addresses these issues:

from sklearn.linear_model import LogisticRegression
from sklearn.metrics import (confusion_matrix, classification_report, 
                             accuracy_score, roc_curve, roc_auc_score)
import matplotlib.pyplot as plt

# Initialize model with custom weights or balanced weights
logmodel = LogisticRegression(solver='liblinear', class_weight={0:0.02, 1:1})
# logmodel = LogisticRegression(solver='liblinear', class_weight='balanced')

# Fit the model ONCE—reuse this trained model for all evaluations
logmodel.fit(X_train, y_train)

# Get hard predictions for confusion matrix/accuracy
predictions = logmodel.predict(X_test)

# Print evaluation metrics
print("Confusion Matrix:")
print(confusion_matrix(y_test, predictions))
print("\nClassification Report:")
print(classification_report(y_test, predictions))
acc = accuracy_score(y_test, predictions)
print(f"Accuracy Score: {acc:.3f}")

# Get positive-class predicted probabilities for ROC
model_probs = logmodel.predict_proba(X_test)
y_score = model_probs[:, 1]  # Extract probabilities for class 1

# Calculate ROC curve metrics and AUC
fpr, tpr, _ = roc_curve(y_test, y_score)
auc = roc_auc_score(y_test, y_score)

# Plot the ROC curve
plt.figure(figsize=(8, 6))
plt.plot(fpr, tpr, marker='.', label=f'ROC Curve (Area = {auc:.2f})')
plt.plot([0, 1], [0, 1], 'r--', label='Random Guess Baseline')
plt.xlabel('False Positive Rate')
plt.ylabel('True Positive Rate')
plt.title('ROC Curve for Weighted Logistic Regression')
plt.legend(loc='lower right')
plt.show()

Quick Notes

  • If you prefer using the decision function instead of predicted probabilities, replace y_score = model_probs[:, 1] with y_score = logmodel.decision_function(X_test)—the ROC/AUC calculation will work the same way.
  • If AUC still doesn't change after trying balanced weights, double-check your data: Are your features actually predictive? Is there extreme class imbalance that's limiting the model's ability to learn?

内容的提问来源于stack exchange,提问作者Arunee Sridee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 18:52:38