GridSearchCV结合EarlyStopping时val_loss不可用的问题咨询
问题原因与解决方案
为什么会报错?
GridSearchCV的交叉验证是用来评估模型最终性能的:它把训练数据拆成多个fold,用训练子集训练模型,验证子集计算指定得分(比如你设置的neg_mean_squared_error)。但Keras的EarlyStopping需要的是训练过程中实时使用的验证数据,这部分数据默认不会由GridSearchCV自动传递给KerasRegressor,所以训练时没有val_loss指标可以监控,直接触发报错。
你提到的validation_split=0.1是让Keras从训练数据中再拆分10%作为验证集,但会缩小实际训练数据的规模,不符合你的需求;而预留的X_test是用来评估最终模型泛化能力的,确实不应该用于超参数优化。
解决方法
核心思路是:让GridSearchCV每个fold中的Keras模型,直接使用该fold的验证子集作为EarlyStopping的监控数据——既不用拆分训练数据,也不会涉及预留的测试集。
由于sklearn原生GridSearchCV不会自动将fold的验证集传递给KerasRegressor的fit方法,我们可以手动实现网格搜索+交叉验证的逻辑:
完整代码
import numpy as np from sklearn.model_selection import KFold, ParameterGrid from sklearn.pipeline import Pipeline from tensorflow.keras.callbacks import EarlyStopping from tensorflow.keras.wrappers.scikit_learn import KerasRegressor # 定义早停回调,明确监控val_loss nn_early_stopping = EarlyStopping( monitor='val_loss', min_delta=0.001, patience=20, restore_best_weights=True ) # 你的模型创建函数(确保输入形状匹配训练数据) def createNN(units=16, activation='relu'): from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Dense model = Sequential() model.add(Dense(units, activation=activation, input_shape=(X_train.shape[1],))) model.add(Dense(1)) # 回归任务输出层无需激活函数 model.compile(optimizer='adam', loss='mse') return model # 参数网格(回归任务不需要softmax,已移除) param_grid = { 'batch_size': [32, 64], 'model__units': [16, 21], 'model__activation': ['relu', 'sigmoid', 'tanh'] } # 初始化3折交叉验证拆分器 kf = KFold(n_splits=3, shuffle=True, random_state=42) # 存储最佳参数和得分 best_score = -np.inf best_params = None # 遍历所有参数组合 for params in ParameterGrid(param_grid): fold_scores = [] # 遍历每个交叉验证fold for train_idx, val_idx in kf.split(X_train): # 拆分当前fold的训练/验证数据 X_train_fold = X_train.iloc[train_idx] y_train_fold = y_train.iloc[train_idx] X_val_fold = X_train.iloc[val_idx] y_val_fold = y_train.iloc[val_idx] # 创建带预处理管道的Keras模型 keras_model = KerasRegressor( model=createNN, epochs=200, callbacks=[nn_early_stopping], batch_size=params['batch_size'], model__units=params['model__units'], model__activation=params['model__activation'] ) pipeline = Pipeline([('preprocessor', preprocessor), ('model', keras_model)]) # 训练时显式传入当前fold的验证集,供早停监控 pipeline.fit( X_train_fold, y_train_fold, model__validation_data=(X_val_fold, y_val_fold) ) # 记录当前fold的得分(和GridSearchCV一致使用负MSE) fold_score = -pipeline.score(X_val_fold, y_val_fold) fold_scores.append(fold_score) # 计算当前参数组合的平均得分 mean_score = np.mean(fold_scores) print(f"参数组合: {params}, 平均负MSE得分: {mean_score:.4f}") # 更新最佳参数 if mean_score > best_score: best_score = mean_score best_params = params print(f"\n最佳参数组合: {best_params}") print(f"最佳平均负MSE得分: {best_score:.4f}") # 用最佳参数训练最终模型 final_keras_model = KerasRegressor( model=createNN, epochs=200, callbacks=[nn_early_stopping], batch_size=best_params['batch_size'], model__units=best_params['model__units'], model__activation=best_params['model__activation'] ) final_pipeline = Pipeline([('preprocessor', preprocessor), ('model', final_keras_model)]) # 若要保留早停,可从X_train拆分极小比例(如5%)作为验证集;若完全不想拆分,注释掉早停回调即可 final_pipeline.fit(X_train, y_train)
关键要点
- 手动遍历参数和fold,确保每个模型训练时都能拿到对应fold的验证集,从根源解决val_loss不可用的问题。
- 回归任务不需要
softmax激活函数,已从参数网格中移除,避免无效参数浪费计算资源。 - 最终模型训练时,若仍想保留早停机制,可拆分极小比例的X_train作为验证集(对训练数据规模影响可忽略);若完全不想拆分训练数据,直接注释掉早停回调即可。
内容的提问来源于stack exchange,提问作者Cheang Wai Bin
相关产品推荐
相关产品推荐

