ROC-AUC、FPR、FNR的Python与R实现及阈值相关技术疑问
Hey there! Let's tackle your questions one by one using the sample data you provided.
roc_curve actual calculated values? Do they match the input scores? Great question! The thresholds returned by scikit-learn's roc_curve are directly derived from your input score values, but there's a small caveat:
- By default,
roc_curveusesdrop_intermediate=True, which removes redundant threshold points that don't change the shape of the ROC curve. This means you might not see every single score value in the returned thresholds. - If you set
drop_intermediate=False, you'll get all unique score values as thresholds, plus an additional-infthreshold (to cover the case where every sample is classified as positive).
Let's verify with your data using Python code:
from sklearn.metrics import roc_curve import numpy as np # Your sample data fraud = np.array([1,1,0,0,0,0,0,0,0,1]) score = np.array([0.84, 1, 1.1, 0.4, 0.6, 0.13, 0.32, 1.4, 0.9, 0.45]) # Get full thresholds (no intermediate drops) fpr, tpr, thresholds = roc_curve(fraud, score, drop_intermediate=False) fnr = 1 - tpr # FNR is 1 minus True Positive Rate print("Thresholds:", thresholds) print("FPR:", fpr.round(4)) print("FNR:", fnr.round(4))
Running this will show you thresholds that include all unique values from your score column, plus the final -inf threshold. The FPR/FNR values here align exactly with manual calculations for each threshold.
To replicate Python's results in R, use the pROC package—it's the standard tool for ROC analysis and lets you control threshold behavior just like scikit-learn. Here's how:
First, install and load the package:
install.packages("pROC") library(pROC)
Then use your sample data to compute the ROC curve and extract metrics:
# Your sample data fraud <- c(1,1,0,0,0,0,0,0,0,1) score <- c(0.84, 1, 1.1, 0.4, 0.6, 0.13, 0.32, 1.4, 0.9, 0.45) # Create ROC object: specify positive class as 1, direction (higher scores = more likely fraud) roc_obj <- roc(response = fraud, predictor = score, pos_label = 1, direction = ">") # Extract all thresholds, FPR, TPR roc_coords <- coords(roc_obj, x = "all", ret = c("threshold", "fpr", "tpr"), transpose = FALSE) roc_coords$fnr <- 1 - roc_coords$tpr # Calculate FNR # Print results (rounded for readability) print(round(roc_coords, 4))
This will give you the exact same thresholds, FPR, and FNR values as Python's drop_intermediate=False setting. The key here is matching the positive label and direction (so R knows higher scores correspond to the fraud class).
Let's cover how to compute these metrics in both languages, both for the full curve and specific thresholds.
Python
ROC-AUC
Use roc_auc_score to get the area under the ROC curve:
from sklearn.metrics import roc_auc_score auc_score = roc_auc_score(fraud, score) print(f"ROC-AUC: {auc_score.round(4)}")
FPR/FNR for a specific threshold
Calculate confusion matrix metrics for a chosen threshold (e.g., 0.5):
from sklearn.metrics import confusion_matrix threshold = 0.5 y_pred = (score >= threshold).astype(int) tn, fp, fn, tp = confusion_matrix(fraud, y_pred).ravel() fpr = fp / (fp + tn) fnr = fn / (fn + tp) print(f"FPR at threshold {threshold}: {fpr.round(4)}") print(f"FNR at threshold {threshold}: {fnr.round(4)}")
R
ROC-AUC
Extract the AUC directly from the roc object:
auc_score <- auc(roc_obj) cat(paste("ROC-AUC:", round(auc_score, 4), "\n"))
FPR/FNR for a specific threshold
Compute confusion matrix metrics manually for a chosen threshold:
threshold <- 0.5 y_pred <- as.integer(score >= threshold) conf_mat <- table(Actual = fraud, Predicted = y_pred) tn <- conf_mat["0", "0"] fp <- conf_mat["0", "1"] fn <- conf_mat["1", "0"] tp <- conf_mat["1", "1"] fpr <- fp / (fp + tn) fnr <- fn / (fn + tp) cat(paste("FPR at threshold", threshold, ":", round(fpr, 4), "\n")) cat(paste("FNR at threshold", threshold, ":", round(fnr, 4), "\n"))
内容的提问来源于stack exchange,提问作者Dino Alessi

