You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在scikit-learn的GridSearchCV中使用需概率的自定义评分函数?

如何让GridSearchCV调用predict_proba()来配合自定义评分函数

要让GridSearchCV使用模型的predict_proba()输出作为自定义评分函数的输入,核心是用sklearn.metrics.make_scorer()包装你的自定义损失函数,并设置needs_proba=True参数——这个参数会明确告诉GridSearchCV:你的评分函数需要的是概率预测值,而非类别标签。

关键步骤说明

  • 用make_scorer包裹自定义函数,添加needs_proba=True:这会触发GridSearchCV在评估时自动调用predict_proba(),把概率结果传入你的自定义函数。
  • 适配概率输出格式:如果是二分类任务,predict_proba()返回的是(n_samples, 2)的矩阵(每行对应样本属于两类的概率),你需要提取对应真实类别的概率列(比如正类概率predictions[:, 1]),否则直接和0/1类型的ground truth做运算会报错;多分类任务则需要根据你的逻辑处理(n_samples, n_classes)的矩阵。

修改后的完整代码

import numpy as np
from sklearn.model_selection import GridSearchCV
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import make_scorer

def my_custom_loss_func(ground_truth, predictions):
    # 针对二分类任务,提取正类的概率(多分类场景请根据需求调整)
    positive_probs = predictions[:, 1]
    diff = np.abs(ground_truth - positive_probs).max()
    return np.log(1 + diff)

# 用make_scorer包装自定义函数,指定需要概率输出
custom_scorer = make_scorer(my_custom_loss_func, needs_proba=True)

# 如果你需要把函数当作损失(值越低越好),可以添加greater_is_better=False
# custom_scorer = make_scorer(my_custom_loss_func, needs_proba=True, greater_is_better=False)

param_grid = {'min_samples_leaf': [1, 2, 5, 10, 20, 50, 100], 'n_estimators': [100, 200, 300]}
grid = GridSearchCV(RandomForestClassifier(), param_grid=param_grid, scoring=custom_scorer)

额外细节提示

  • greater_is_better参数:默认情况下make_scorer会把你的函数视为得分函数(值越高模型越好),如果你的自定义函数是损失函数(值越低模型越好),一定要加上greater_is_better=False,这样GridSearchCV才会正确寻找最优参数组合。

内容的提问来源于stack exchange,提问作者mgoldwasser

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 04:08:08