逻辑回归负类Precision-Recall Curve精确率与召回率正相关问题排查
问题原因及解决方法
问题根源
precision_recall_curve 默认将真实标签(y_true)中的1视为正类,且要求传入的 probas_pred 是正类的预测概率。当你分析类别0的PR曲线时:
- 你传入了类别0的预测概率,但未将真实标签转换为“类别0为正类”的二值标签
- 此时函数仍以原标签中的1作为正类,而类别0的概率越高,模型越倾向于判断为负类(原标签0),导致PR曲线的逻辑完全反转,出现精确率随召回率上升的异常情况
修正后的代码
把真实标签转换为以类别0为正类的形式,同时保留类别0的预测概率作为得分:
from sklearn import linear_model from sklearn.metrics import precision_recall_fscore_support, precision_recall_curve import matplotlib.pyplot as plt clf = linear_model.LogisticRegression().fit(ohe_X_train, y_train) clf_predictions = clf.predict(ohe_X_test) class_of_interest = 0 # 计算类别0的P/R指标 precision, recall, fscore, support = precision_recall_fscore_support(y_test, clf_predictions, labels=[class_of_interest]) # 修正PR曲线计算:将真实标签转换为"是否为类别0"的二值标签 y_true_interest = (y_test == class_of_interest).astype(int) y_scores_interest = clf.predict_proba(ohe_X_test)[:, class_of_interest] precision_curve, recall_curve, thresholds = precision_recall_curve(y_true_interest, y_scores_interest) plt.plot(recall_curve, precision_curve, marker='.') plt.xlabel('Recall') plt.ylabel('Precision') plt.title('Precision-Recall Curve for Class 0') plt.grid(True) plt.show()
额外说明
- 当分析类别1时无需转换,因为默认正类就是1,所以曲线表现正常
- 转换后的
y_true_interest中,1代表样本属于类别0(正类),0代表不属于(负类),此时传入的类别0预测概率与正类定义匹配,PR曲线的计算逻辑就正确了
内容的提问来源于stack exchange,提问作者oliver13
相关产品推荐
相关产品推荐

