You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何进一步拆分hold out数据集以拟合校准模型?

How to Split Your Holdout Dataset for Model Calibration

Got it, let's break this down clearly—you need a separate dataset to fit your calibration model without messing up your final evaluation, right? Here's a step-by-step approach tailored to your existing data setup:

Why Split the Test Set?

Model calibration adjusts the predicted probabilities of your trained model to be more accurate (e.g., a predicted 80% fraud probability should actually correspond to an 80% real-world fraud rate). To do this properly:

  • You can't use your original training set (calibrating on training data will overfit the calibration logic).
  • You shouldn't use your full test set (this would bias your final performance evaluation).

So we'll split your existing Test dataset into two distinct parts: a calibration set (to tune the calibration) and a final test set (to evaluate the calibrated model's real-world performance).

Step 1: Split the Holdout Dataset

Use stratified splitting to preserve the class distribution (critical for imbalanced fraud detection data). Here's the code:

import pandas as pd
from sklearn.model_selection import train_test_split

# Your existing test data setup (just confirming)
Test = pd.read_csv("Test.csv")
ytest = Test['Fraud']
Xtest = Test.drop(['Fraud'], axis=1)

# Split into calibration set (80%) and final test set (20%)
X_calib, X_final_test, y_calib, y_final_test = train_test_split(
    Xtest, ytest,
    test_size=0.2,  # Adjust this ratio based on your dataset size
    random_state=42,  # Fixed seed for reproducibility
    stratify=ytest  # Keeps fraud/non-fraud ratio consistent across splits
)

Step 2: Train Your Base Model + Fit Calibration

First, train your core model using your original training data (Xtrain/ytrain). Then use the calibration set to adjust the probabilities. Here's an example with a random forest base model and Platt scaling (sigmoid calibration):

from sklearn.ensemble import RandomForestClassifier
from sklearn.calibration import CalibratedClassifierCV

# Train your base prediction model
base_model = RandomForestClassifier(random_state=42)
base_model.fit(Xtrain, ytrain)

# Fit the calibration model using the calibration set
# Use `cv='prefit'` because we already trained the base model
calibrated_model = CalibratedClassifierCV(
    base_model,
    method='sigmoid',  # Use 'isotonic' if you have a large calibration set
    cv='prefit'
)
calibrated_model.fit(X_calib, y_calib)

Quick notes on calibration methods:

  • sigmoid (Platt scaling) works well for models with roughly log-odds outputs (like logistic regression, SVMs).
  • isotonic is more flexible but requires a larger calibration set to avoid overfitting.

Step 3: Evaluate the Calibrated Model

Now use the final test set to assess how well your calibrated model performs—this gives you an unbiased measure of both predictive accuracy and probability calibration (use metrics like Brier score, reliability plots, or precision-recall curves):

from sklearn.metrics import brier_score_loss

# Get calibrated probabilities for the fraud class
calibrated_probs = calibrated_model.predict_proba(X_final_test)[:, 1]

# Calculate Brier score (lower value = better calibration)
brier_score = brier_score_loss(y_final_test, calibrated_probs)
print(f"Calibrated Model Brier Score: {brier_score:.4f}")

Key Tips

  • If your test dataset is very small, skip splitting it and use cross-validation during calibration (set cv=5 in CalibratedClassifierCV instead of cv='prefit'—but make sure your training set is large enough).
  • Always keep the final test set untouched until the very end—don't use it for any tuning or calibration steps.

内容的提问来源于stack exchange,提问作者Tumi Sebela

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 13:33:10