You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用CalibratedClassifier包装随机森林后,如何评估训练集性能判断过拟合?

How to Evaluate Training Set Performance for Calibrated Random Forests (and Detect Overfitting)

Great question—this is a super common pain point when combining ensemble models like Random Forests with probability calibration in scikit-learn. The key insight here is that calibration only adjusts the scaling of probability outputs, not the model's core predictive ability or decision boundaries. That means we can leverage the original uncalibrated Random Forest to assess overfitting, while still using the calibrated model for its improved probability estimates on new data.

Here are the most reliable approaches:

1. Use the Original Uncalibrated Random Forest's OOB Metrics

Random Forest's oob_decision_function_ and oob_score_ are still your best tools for evaluating out-of-bag (unseen during training) performance, which directly signals overfitting. Since calibration doesn't change how the model ranks samples, these metrics are fully representative of the calibrated model's generalization ability.

Example Code:

from sklearn.ensemble import RandomForestClassifier
from sklearn.calibration import CalibratedClassifierCV
from sklearn.datasets import make_classification
from sklearn.metrics import roc_auc_score

# Generate sample data
X, y = make_classification(n_samples=1000, n_classes=2, random_state=42)

# Train uncalibrated Random Forest with OOB enabled
rf_uncal = RandomForestClassifier(
    oob_score=True,
    oob_decision_function=True,
    random_state=42
)
rf_uncal.fit(X, y)

# Calculate OOB AUC (measures generalization ability)
oob_probs = rf_uncal.oob_decision_function_[:, 1]
oob_auc = roc_auc_score(y, oob_probs)
print(f"OOB AUC (uncalibrated): {oob_auc:.3f}")

# Compare to training set AUC (detect overfitting)
train_probs_uncal = rf_uncal.predict_proba(X)[:, 1]
train_auc_uncal = roc_auc_score(y, train_probs_uncal)
print(f"Training Set AUC (uncalibrated): {train_auc_uncal:.3f}")

# If training AUC is drastically higher than OOB AUC, you have overfitting

# Now calibrate the model for reliable probability outputs
calibrated_rf = CalibratedClassifierCV(
    rf_uncal,
    method="sigmoid",  # Use "isotonic" for non-monotonic data
    cv=5
)
calibrated_rf.fit(X, y)

2. Split Training Data into Training + Calibration Subsets

If you want to directly evaluate the calibrated model's behavior without relying on OOB, split your original training data into two parts:

  • A training subset to fit the base Random Forest
  • A calibration subset to calibrate the probabilities

This lets you compare the base model's performance on the training subset (how well it fits the data) to the calibrated model's performance on the calibration subset (how well it generalizes), which is a clear signal of overfitting.

Example Code:

from sklearn.model_selection import train_test_split
from sklearn.metrics import brier_score_loss

# Split original training data
X_train, X_cal, y_train, y_cal = train_test_split(
    X, y, test_size=0.2, random_state=42
)

# Train base Random Forest on the training subset
rf_base = RandomForestClassifier(random_state=42)
rf_base.fit(X_train, y_train)

# Calibrate using the calibration subset (use cv="prefit" since model is already trained)
calibrated_rf = CalibratedClassifierCV(
    rf_base,
    method="sigmoid",
    cv="prefit"
)
calibrated_rf.fit(X_cal, y_cal)

# Evaluate overfitting: compare training vs calibration performance
train_acc = (rf_base.predict(X_train) == y_train).mean()
cal_acc = (calibrated_rf.predict(X_cal) == y_cal).mean()

print(f"Training Subset Accuracy: {train_acc:.3f}")
print(f"Calibration Subset Accuracy: {cal_acc:.3f}")

# Check calibration quality with Brier score (lower = better)
cal_brier = brier_score_loss(y_cal, calibrated_rf.predict_proba(X_cal)[:, 1])
print(f"Calibration Subset Brier Score (calibrated): {cal_brier:.3f}")

3. Use Calibrated Model's Class Predictions (Not Probabilities) on Training Data

While you shouldn't trust the calibrated model's predict_proba outputs on the training set (risk of data leakage from calibration), you can still use its predict method to get class labels. Comparing training set accuracy to validation set accuracy will tell you if the model is overfitting—since calibration doesn't change the core classification logic (just probability scaling), these accuracy metrics are valid.

Key Notes:

  • Never use calibrated predict_proba on training data: Calibration is fit on held-out data, so applying it to the training set will produce biased, over-optimistic probability estimates.
  • Calibration doesn't fix overfitting: If your base Random Forest is overfitting, calibration won't fix it—it just makes the overfit model's probabilities more reliable. Always address overfitting first (e.g., reduce tree depth, increase min_samples_split) before calibrating.

内容的提问来源于stack exchange,提问作者pipefish

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:27:13