scikit-learn是否有计算二分类器全误差曲线或多指标的便捷方法?
scikit-learn中一次性获取多分类曲线指标的方法
scikit-learn目前没有公开的内置函数能一次性返回TP/TN/FP/FN及对应阈值,也没有直接一次性计算ROC、DET、PR曲线所有指标的方法,但可以通过以下两种方式实现你的需求:
方法1:基于内置函数推导指标
利用roc_curve返回的FPR、TPR和阈值,结合真实正负样本数量,就能直接推导TP、TN、FP、FN,这种方式完全沿用scikit-learn的计算逻辑,几乎不会有误差:
from sklearn.metrics import roc_curve import numpy as np y_true = np.array([0, 1, 0, 1, 1, 0]) y_score = np.array([0.1, 0.8, 0.3, 0.9, 0.7, 0.2]) # 获取ROC曲线的核心指标 fpr, tpr, thresholds = roc_curve(y_true, y_score) # 统计真实正负样本数 pos_count = np.sum(y_true) neg_count = len(y_true) - pos_count # 推导各类计数 tp = tpr * pos_count fp = fpr * neg_count tn = neg_count - fp fn = pos_count - tp
得到这些计数后,你可以轻松计算精确率、召回率等任意指标。
方法2:自定义函数直接返回计数与阈值
如果你想要一个类似errors_curve的函数,可以自己实现,逻辑和scikit-learn内置函数对齐:
import numpy as np def errors_curve(y_true, y_score): # 按预测得分降序排序样本 sorted_indices = np.argsort(y_score)[::-1] y_true_sorted = y_true[sorted_indices] y_score_sorted = y_score[sorted_indices] pos_count = np.sum(y_true) neg_count = len(y_true) - pos_count # 计算累计TP和FP tp = np.cumsum(y_true_sorted) fp = np.cumsum(1 - y_true_sorted) # 推导TN和FN tn = neg_count - fp fn = pos_count - tp # 生成阈值:取相邻得分的中点,最后补充一个高于最大得分的阈值 thresholds = np.concatenate([ np.maximum(y_score_sorted[1:], y_score_sorted[:-1]), [y_score_sorted[0] + 1e-6] ]) return fp, tp, fn, tn, thresholds
调用这个函数后,就能直接得到你需要的所有值,后续可基于此计算任意曲线指标。
关于一次性获取所有曲线指标
目前scikit-learn没有现成函数能一次性输出ROC、DET、PR曲线的全部指标,但你可以基于上述方法得到的计数,自行计算各类指标,或者分别调用roc_curve、det_curve、precision_recall_curve三个函数——两种方式的计算结果完全一致。
内容的提问来源于stack exchange,提问作者william_grisaitis
相关产品推荐
相关产品推荐

