You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

scikit-learn是否有计算二分类器全误差曲线或多指标的便捷方法?

scikit-learn中一次性获取多分类曲线指标的方法

scikit-learn目前没有公开的内置函数能一次性返回TP/TN/FP/FN及对应阈值,也没有直接一次性计算ROC、DET、PR曲线所有指标的方法,但可以通过以下两种方式实现你的需求:

方法1:基于内置函数推导指标

利用roc_curve返回的FPR、TPR和阈值,结合真实正负样本数量,就能直接推导TP、TN、FP、FN,这种方式完全沿用scikit-learn的计算逻辑,几乎不会有误差:

from sklearn.metrics import roc_curve
import numpy as np

y_true = np.array([0, 1, 0, 1, 1, 0])
y_score = np.array([0.1, 0.8, 0.3, 0.9, 0.7, 0.2])

# 获取ROC曲线的核心指标
fpr, tpr, thresholds = roc_curve(y_true, y_score)
# 统计真实正负样本数
pos_count = np.sum(y_true)
neg_count = len(y_true) - pos_count

# 推导各类计数
tp = tpr * pos_count
fp = fpr * neg_count
tn = neg_count - fp
fn = pos_count - tp

得到这些计数后,你可以轻松计算精确率、召回率等任意指标。

方法2:自定义函数直接返回计数与阈值

如果你想要一个类似errors_curve的函数,可以自己实现,逻辑和scikit-learn内置函数对齐:

import numpy as np

def errors_curve(y_true, y_score):
    # 按预测得分降序排序样本
    sorted_indices = np.argsort(y_score)[::-1]
    y_true_sorted = y_true[sorted_indices]
    y_score_sorted = y_score[sorted_indices]
    
    pos_count = np.sum(y_true)
    neg_count = len(y_true) - pos_count
    
    # 计算累计TP和FP
    tp = np.cumsum(y_true_sorted)
    fp = np.cumsum(1 - y_true_sorted)
    # 推导TN和FN
    tn = neg_count - fp
    fn = pos_count - tp
    
    # 生成阈值:取相邻得分的中点,最后补充一个高于最大得分的阈值
    thresholds = np.concatenate([
        np.maximum(y_score_sorted[1:], y_score_sorted[:-1]),
        [y_score_sorted[0] + 1e-6]
    ])
    
    return fp, tp, fn, tn, thresholds

调用这个函数后,就能直接得到你需要的所有值,后续可基于此计算任意曲线指标。

关于一次性获取所有曲线指标

目前scikit-learn没有现成函数能一次性输出ROC、DET、PR曲线的全部指标,但你可以基于上述方法得到的计数,自行计算各类指标,或者分别调用roc_curve、det_curve、precision_recall_curve三个函数——两种方式的计算结果完全一致。

内容的提问来源于stack exchange,提问作者william_grisaitis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 12:23:17