如何获取GridSearchCV特定折的数据点及测试样本预测值?
解决GridSearchCV交叉验证中特定折样本与预测值获取问题
一、定位特定交叉验证折的样本数据
通过实例化带固定随机状态的KFold对象,我们可以提前保存所有折的拆分索引,后续直接用索引定位对应样本:
import numpy as np from sklearn.model_selection import KFold from sklearn.pipeline import Pipeline from sklearn.preprocessing import MinMaxScaler from sklearn.ensemble import GradientBoostingRegressor from sklearn.model_selection import GridSearchCV # 超参数网格(保持你提供的命名与结构) XGBoost_param_grid = { "learning_rate": [0.001, 0.01, 0.1, 0.5], "n_estimators": [10, 100, 500, 1000], "max_depth": [3, 10, None], "max_features": ["sqrt", "log2", None] } # 初始化KFold并保存拆分索引(固定random_state保证拆分一致) kf = KFold(n_splits=5, shuffle=True, random_state=42) split_indices = list(kf.split(X)) # 每个元素是(train_idx, test_idx)元组 # 传入KFold到GridSearchCV gs_obj = GridSearchCV( estimator=GradientBoostingRegressor(), param_grid=XGBoost_param_grid, cv=kf ) pipeline = Pipeline([("scaler", MinMaxScaler()), ("model", gs_obj)]) pipeline.fit(X, y) # 获取第4折的测试样本(索引从0开始) _, fold4_test_idx = split_indices[4] # 若X是DataFrame用iloc,numpy数组直接用X[fold4_test_idx] fold4_test_samples = X.iloc[fold4_test_idx] fold4_test_labels = y.iloc[fold4_test_idx]
二、获取测试折中每个样本的预测值
GridSearchCV默认不保存样本级预测值,可通过以下两种方式获取:
方法1:基于最优模型重新遍历各折预测
从Pipeline中取出最优模型,对每个折的训练数据重新拟合Pipeline后预测测试样本:
# 获取训练好的最优模型 best_model = pipeline.named_steps['model'].best_estimator_ # 遍历所有折,记录预测值 fold_pred_details = [] for train_idx, test_idx in split_indices: X_train, X_test = X.iloc[train_idx], X.iloc[test_idx] y_train = y.iloc[train_idx] # 拟合当前折的Pipeline(保证scaler是基于当前折训练数据拟合的) fold_pipeline = Pipeline([("scaler", MinMaxScaler()), ("model", best_model)]) fold_pipeline.fit(X_train, y_train) # 保存当前折的索引、真实值和预测值 fold_pred_details.append({ "test_indices": test_idx, "true_labels": y.iloc[test_idx].values, "predictions": fold_pipeline.predict(X_test) }) # 提取第4折的预测结果 fold4_predictions = fold_pred_details[4]
方法2:自定义评分函数记录预测值
通过cross_validate配合自定义评分函数,在交叉验证过程中直接记录预测值,避免重复训练:
from sklearn.model_selection import cross_validate from sklearn.metrics import r2_score # 用于保存预测信息的字典 pred_log = {"true_labels": [], "predictions": []} def score_with_pred_log(model, X_test, y_test): y_pred = model.predict(X_test) pred_log["true_labels"].append(y_test) pred_log["predictions"].append(y_pred) return r2_score(y_test, y_pred) # 取出最优参数构建最终模型 best_params = pipeline.named_steps['model'].best_params_ final_pipeline = Pipeline([ ("scaler", MinMaxScaler()), ("model", GradientBoostingRegressor(**best_params)) ]) # 运行交叉验证并记录预测值 cross_validate( final_pipeline, X, y, cv=kf, scoring={"r2": score_with_pred_log}, return_estimator=True ) # pred_log中已保存所有折的真实值和预测值,对应split_indices的顺序
关键注意点
- 必须固定
KFold的random_state和shuffle参数,确保拆分索引与GridSearchCV使用的完全一致,避免样本匹配错误。 - 若数据为numpy数组,直接使用
X[test_idx]替代X.iloc[test_idx]即可。
内容的提问来源于stack exchange,提问作者Gojih
相关产品推荐
相关产品推荐

