You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于make_scorer与GridSearchCV的技术疑问及自定义指标实现

Hey there! Let's break down your questions one by one, with clear examples tied to your code snippets.

1. Do I need to pass (y_true, y_pred) to make_scorer? If so, how? Can you give an example?

No, you don’t manually pass y_true and y_pred directly to make_scorer. Instead, make_scorer wraps a custom scoring function that expects y_true and y_pred as its first two arguments. When you use this scorer with tools like GridSearchCV, the cross-validation process automatically feeds the correct true and predicted values from each fold into your function behind the scenes.

Here’s how to adapt your custom z-value calculation into a function that works seamlessly with make_scorer:

import numpy as np
from sklearn.metrics import make_scorer

def custom_z_score(y_true, y_pred):
    # Calculate r: number of true negatives (predicted 0 and matches ground truth)
    r = np.sum((y_pred == 0) & (y_pred == y_true))
    # Calculate s: number of false positives (predicted 1 but doesn't match ground truth)
    s = np.sum((y_pred == 1) & (y_pred != y_true))
    # Avoid division by zero edge case
    if s == 0:
        return 0.0  # Or another appropriate default value for your use case
    return r / s

# Wrap your custom function into a scorer object
z_scorer = make_scorer(custom_z_score)

You don’t need to pass y_true/y_pred to make_scorer here—those values are handled automatically during cross-validation.

2. How to set a custom evaluation criterion in the scoring parameter?

Once you’ve created your scorer object with make_scorer, simply pass it directly to the scoring parameter of GridSearchCV (or other sklearn tools like cross_val_score). Using your existing code as a base, here’s how to integrate the custom z-score:

from sklearn.model_selection import GridSearchCV
# Assume clf (your classifier) and parameter_grid are already defined
grid_searcher = GridSearchCV(clf, parameter_grid, verbose=200, scoring=z_scorer)
grid_searcher.fit(X_train, y_train)
clf_best = grid_searcher.best_estimator_

If you want to use multiple scoring metrics (including both custom and built-in ones), you can pass a dictionary where keys are metric names and values are scorer objects:

from sklearn.metrics import f1_score

# Create a dictionary of multiple scorers
scorers = {
    'z_score': z_scorer,
    'f1_class_0': make_scorer(f1_score, pos_label=0)
}

# Pass the dictionary to scoring, and specify which metric to use for refitting
grid_searcher = GridSearchCV(clf, parameter_grid, verbose=200, scoring=scorers, refit='z_score')

The refit parameter lets you choose which metric to prioritize when selecting the best-performing model.

3. Are the iteration results (scores in CV output) from the training set or test set?

The scores you see in the CV iteration output (like score=0.4419706300331596) are from the validation fold of the cross-validation split.

When using GridSearchCV, your training data (X_train, y_train) is split into k folds. For each iteration:

  • The model is trained on k-1 of those folds (the "training" portion of the split)
  • The score is calculated on the remaining 1 fold (the held-out validation set)

This means the scores are not from your full training set, nor from an external test set—they’re from the validation subset held out during each cross-validation run. This helps you estimate generalization performance without touching your test data.


内容的提问来源于stack exchange,提问作者user287629

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:21:53