You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GridSearchCV结合EarlyStopping时val_loss不可用的问题咨询

问题原因与解决方案

为什么会报错?

GridSearchCV的交叉验证是用来评估模型最终性能的:它把训练数据拆成多个fold,用训练子集训练模型,验证子集计算指定得分(比如你设置的neg_mean_squared_error)。但Keras的EarlyStopping需要的是训练过程中实时使用的验证数据,这部分数据默认不会由GridSearchCV自动传递给KerasRegressor,所以训练时没有val_loss指标可以监控,直接触发报错。

你提到的validation_split=0.1是让Keras从训练数据中再拆分10%作为验证集,但会缩小实际训练数据的规模,不符合你的需求;而预留的X_test是用来评估最终模型泛化能力的,确实不应该用于超参数优化。

解决方法

核心思路是:让GridSearchCV每个fold中的Keras模型,直接使用该fold的验证子集作为EarlyStopping的监控数据——既不用拆分训练数据,也不会涉及预留的测试集。

由于sklearn原生GridSearchCV不会自动将fold的验证集传递给KerasRegressor的fit方法,我们可以手动实现网格搜索+交叉验证的逻辑:

完整代码

import numpy as np
from sklearn.model_selection import KFold, ParameterGrid
from sklearn.pipeline import Pipeline
from tensorflow.keras.callbacks import EarlyStopping
from tensorflow.keras.wrappers.scikit_learn import KerasRegressor

# 定义早停回调,明确监控val_loss
nn_early_stopping = EarlyStopping(
    monitor='val_loss',
    min_delta=0.001,
    patience=20,
    restore_best_weights=True
)

# 你的模型创建函数(确保输入形状匹配训练数据)
def createNN(units=16, activation='relu'):
    from tensorflow.keras.models import Sequential
    from tensorflow.keras.layers import Dense
    model = Sequential()
    model.add(Dense(units, activation=activation, input_shape=(X_train.shape[1],)))
    model.add(Dense(1))  # 回归任务输出层无需激活函数
    model.compile(optimizer='adam', loss='mse')
    return model

# 参数网格(回归任务不需要softmax,已移除)
param_grid = {
    'batch_size': [32, 64],
    'model__units': [16, 21],
    'model__activation': ['relu', 'sigmoid', 'tanh']
}

# 初始化3折交叉验证拆分器
kf = KFold(n_splits=3, shuffle=True, random_state=42)

# 存储最佳参数和得分
best_score = -np.inf
best_params = None

# 遍历所有参数组合
for params in ParameterGrid(param_grid):
    fold_scores = []
    # 遍历每个交叉验证fold
    for train_idx, val_idx in kf.split(X_train):
        # 拆分当前fold的训练/验证数据
        X_train_fold = X_train.iloc[train_idx]
        y_train_fold = y_train.iloc[train_idx]
        X_val_fold = X_train.iloc[val_idx]
        y_val_fold = y_train.iloc[val_idx]
        
        # 创建带预处理管道的Keras模型
        keras_model = KerasRegressor(
            model=createNN,
            epochs=200,
            callbacks=[nn_early_stopping],
            batch_size=params['batch_size'],
            model__units=params['model__units'],
            model__activation=params['model__activation']
        )
        pipeline = Pipeline([('preprocessor', preprocessor), ('model', keras_model)])
        
        # 训练时显式传入当前fold的验证集,供早停监控
        pipeline.fit(
            X_train_fold, y_train_fold,
            model__validation_data=(X_val_fold, y_val_fold)
        )
        
        # 记录当前fold的得分(和GridSearchCV一致使用负MSE)
        fold_score = -pipeline.score(X_val_fold, y_val_fold)
        fold_scores.append(fold_score)
    
    # 计算当前参数组合的平均得分
    mean_score = np.mean(fold_scores)
    print(f"参数组合: {params}, 平均负MSE得分: {mean_score:.4f}")
    
    # 更新最佳参数
    if mean_score > best_score:
        best_score = mean_score
        best_params = params

print(f"\n最佳参数组合: {best_params}")
print(f"最佳平均负MSE得分: {best_score:.4f}")

# 用最佳参数训练最终模型
final_keras_model = KerasRegressor(
    model=createNN,
    epochs=200,
    callbacks=[nn_early_stopping],
    batch_size=best_params['batch_size'],
    model__units=best_params['model__units'],
    model__activation=best_params['model__activation']
)
final_pipeline = Pipeline([('preprocessor', preprocessor), ('model', final_keras_model)])

# 若要保留早停,可从X_train拆分极小比例(如5%)作为验证集;若完全不想拆分,注释掉早停回调即可
final_pipeline.fit(X_train, y_train)

关键要点

  • 手动遍历参数和fold,确保每个模型训练时都能拿到对应fold的验证集,从根源解决val_loss不可用的问题。
  • 回归任务不需要softmax激活函数,已从参数网格中移除,避免无效参数浪费计算资源。
  • 最终模型训练时,若仍想保留早停机制,可拆分极小比例的X_train作为验证集(对训练数据规模影响可忽略);若完全不想拆分训练数据,直接注释掉早停回调即可。

内容的提问来源于stack exchange,提问作者Cheang Wai Bin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 11:50:24