You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取GridSearchCV特定折的数据点及测试样本预测值?

解决GridSearchCV交叉验证中特定折样本与预测值获取问题

一、定位特定交叉验证折的样本数据

通过实例化带固定随机状态的KFold对象,我们可以提前保存所有折的拆分索引,后续直接用索引定位对应样本:

import numpy as np
from sklearn.model_selection import KFold
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import MinMaxScaler
from sklearn.ensemble import GradientBoostingRegressor
from sklearn.model_selection import GridSearchCV

# 超参数网格(保持你提供的命名与结构)
XGBoost_param_grid = {
    "learning_rate": [0.001, 0.01, 0.1, 0.5],
    "n_estimators": [10, 100, 500, 1000],
    "max_depth": [3, 10, None],
    "max_features": ["sqrt", "log2", None]
}

# 初始化KFold并保存拆分索引(固定random_state保证拆分一致)
kf = KFold(n_splits=5, shuffle=True, random_state=42)
split_indices = list(kf.split(X))  # 每个元素是(train_idx, test_idx)元组

# 传入KFold到GridSearchCV
gs_obj = GridSearchCV(
    estimator=GradientBoostingRegressor(),
    param_grid=XGBoost_param_grid,
    cv=kf
)

pipeline = Pipeline([("scaler", MinMaxScaler()), ("model", gs_obj)])
pipeline.fit(X, y)

# 获取第4折的测试样本(索引从0开始)
_, fold4_test_idx = split_indices[4]
# 若X是DataFrame用iloc,numpy数组直接用X[fold4_test_idx]
fold4_test_samples = X.iloc[fold4_test_idx]
fold4_test_labels = y.iloc[fold4_test_idx]

二、获取测试折中每个样本的预测值

GridSearchCV默认不保存样本级预测值,可通过以下两种方式获取:

方法1:基于最优模型重新遍历各折预测

从Pipeline中取出最优模型,对每个折的训练数据重新拟合Pipeline后预测测试样本:

# 获取训练好的最优模型
best_model = pipeline.named_steps['model'].best_estimator_

# 遍历所有折,记录预测值
fold_pred_details = []
for train_idx, test_idx in split_indices:
    X_train, X_test = X.iloc[train_idx], X.iloc[test_idx]
    y_train = y.iloc[train_idx]
    
    # 拟合当前折的Pipeline(保证scaler是基于当前折训练数据拟合的)
    fold_pipeline = Pipeline([("scaler", MinMaxScaler()), ("model", best_model)])
    fold_pipeline.fit(X_train, y_train)
    
    # 保存当前折的索引、真实值和预测值
    fold_pred_details.append({
        "test_indices": test_idx,
        "true_labels": y.iloc[test_idx].values,
        "predictions": fold_pipeline.predict(X_test)
    })

# 提取第4折的预测结果
fold4_predictions = fold_pred_details[4]

方法2:自定义评分函数记录预测值

通过cross_validate配合自定义评分函数,在交叉验证过程中直接记录预测值,避免重复训练:

from sklearn.model_selection import cross_validate
from sklearn.metrics import r2_score

# 用于保存预测信息的字典
pred_log = {"true_labels": [], "predictions": []}

def score_with_pred_log(model, X_test, y_test):
    y_pred = model.predict(X_test)
    pred_log["true_labels"].append(y_test)
    pred_log["predictions"].append(y_pred)
    return r2_score(y_test, y_pred)

# 取出最优参数构建最终模型
best_params = pipeline.named_steps['model'].best_params_
final_pipeline = Pipeline([
    ("scaler", MinMaxScaler()),
    ("model", GradientBoostingRegressor(**best_params))
])

# 运行交叉验证并记录预测值
cross_validate(
    final_pipeline,
    X, y,
    cv=kf,
    scoring={"r2": score_with_pred_log},
    return_estimator=True
)

# pred_log中已保存所有折的真实值和预测值,对应split_indices的顺序

关键注意点

  • 必须固定KFold的random_state和shuffle参数,确保拆分索引与GridSearchCV使用的完全一致,避免样本匹配错误。
  • 若数据为numpy数组,直接使用X[test_idx]替代X.iloc[test_idx]即可。

内容的提问来源于stack exchange,提问作者Gojih

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 15:06:08