You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为CalibratedClassifierCV复现CatBoost指定FPR阈值选择方法

解决方法

核心思路是找到FPR≤0.1时对应的最大分类阈值——阈值越高,模型判定为正类的样本越少,FPR也就越低,这和CatBoost的select_threshold逻辑一致。具体实现如下:

修改后的代码:

import numpy as np
from sklearn import metrics

prob_pred = model.predict_proba(X[features_list])[:, 1]
            
fpr, tpr, thresholds = metrics.roc_curve(X['target'], prob_pred)

# 指定目标FPR值
target_fpr = 0.1

# 筛选出所有FPR不超过目标值的索引
valid_indices = np.where(fpr <= target_fpr)[0]
# 取最后一个索引(对应满足条件的最大阈值,因为thresholds随FPR升高递减)
optimal_idx = valid_indices[-1]
boundary = thresholds[optimal_idx]

# 可选:如果需要更精确的阈值(无刚好匹配0.1的FPR时),用线性插值计算
# if fpr[-1] < target_fpr:
#     boundary = thresholds[-1]  # 所有样本都满足FPR要求,取最小阈值(全判正)
# else:
#     idx = np.argmax(fpr > target_fpr)
#     # 用前后两个点插值计算对应阈值
#     fpr_prev, fpr_curr = fpr[idx-1], fpr[idx]
#     thresh_prev, thresh_curr = thresholds[idx-1], thresholds[idx]
#     boundary = thresh_prev + (thresh_curr - thresh_prev) * (target_fpr - fpr_prev) / (fpr_curr - fpr_prev)

binary_pred = [1 if i >= boundary else 0 for i in prob_pred]

关键说明

  • sklearn.metrics.roc_curve返回的thresholds是递减数组:阈值越低,FPR和TPR同步升高
  • 直接取FPR<=0.1的最后一个索引,能得到满足条件的最高阈值,避免不必要的误判
  • 注释部分的插值逻辑,可在无刚好匹配目标FPR的点时,计算更精准的阈值

内容的提问来源于stack exchange,提问作者Michael

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 20:39:26