基于Keras Functional API的多输入回归模型超参数调优求助
解决Keras Functional多输入模型的超参数调优问题
我明白你现在的困境——用Functional API搭建了多输入回归模型,但GridSearchCV和Hyperas没法直接支持,单独调分支又不合理。其实有几个实用的方案能搞定这个问题,下面给你详细说:
1. 用KerasRegressor包装模型适配GridSearchCV
GridSearchCV需要符合scikit-learn的estimator接口,所以我们可以把Functional模型用KerasRegressor包装起来,关键是要把超参数作为参数传入模型构建函数,而不是像你原来那样把输入数据传进去(原来的build_model设计有问题,模型构建不该依赖具体数据)。
修改模型构建函数
先重构你的build_model,让它接受超参数,输入层的shape从外部传入:
from tensorflow.keras.models import Model from tensorflow.keras.layers import Input, Dense, BatchNormalization, Concatenate, Dropout from tensorflow.keras.regularizers import l2 from tensorflow.keras.optimizers import AdaBound def build_model(input_shape1, input_shape2, units1=5, units2=64, dropout_rate=0.08, lr=0.001, final_lr=0.1): """ 创建多输入ANN模型,接受超参数和输入形状 :param input_shape1: 第一个输入的形状(不含batch维度) :param input_shape2: 第二个输入的形状(不含batch维度) :param units1: 第一个分支Dense层的单元数 :param units2: 第二个分支Dense层的单元数 :param dropout_rate: Dropout层的丢弃率 :param lr: 优化器初始学习率 :param final_lr: AdaBound的最终学习率 :return: 编译好的模型 """ # 定义输入层 input1 = Input(shape=input_shape1, name="Input1") input2 = Input(shape=input_shape2, name="Input2") # 第一个分支 x = BatchNormalization()(input1) x = Dense(units=units1, activation='relu', kernel_regularizer=l2(0.001))(x) # 第二个分支 y = BatchNormalization()(input2) y = Dense(units=units2, activation="relu", kernel_regularizer=l2(0.001))(y) # 合并分支 concatenated = Concatenate()([x, y]) # 输出层 outputs = Dropout(dropout_rate)(concatenated) outputs = Dense(1, name="output")(outputs) # 创建并编译模型 model = Model(inputs=[input1, input2], outputs=outputs) model.compile(loss='mse', optimizer=AdaBound(lr=lr, final_lr=final_lr), metrics=['mse']) return model
用KerasRegressor包装并做GridSearchCV
然后定义一个包装函数,把超参数传递给build_model,再用GridSearchCV搜索:
from tensorflow.keras.wrappers.scikit_learn import KerasRegressor from sklearn.model_selection import GridSearchCV import numpy as np # 假设你的输入数据是X1, X2,标签是y # 先处理数据形状(这里假设你已经做好了预处理) input_shape1 = X1.shape[1:] input_shape2 = X2.shape[1:] # 定义包装函数 def create_model(input_shape1, input_shape2): def wrapper(units1, units2, dropout_rate, lr, final_lr): return build_model(input_shape1, input_shape2, units1=units1, units2=units2, dropout_rate=dropout_rate, lr=lr, final_lr=final_lr) return wrapper # 初始化KerasRegressor model = KerasRegressor(build_fn=create_model(input_shape1, input_shape2), verbose=1) # 设置超参数搜索空间 param_grid = { 'units1': [5, 10, 16], 'units2': [32, 64, 128], 'dropout_rate': [0.05, 0.08, 0.1], 'lr': [0.0001, 0.001, 0.01], 'final_lr': [0.05, 0.1, 0.2] } # 运行GridSearchCV grid = GridSearchCV(estimator=model, param_grid=param_grid, cv=3, scoring='neg_mean_squared_error') grid_result = grid.fit([X1, X2], y) # 输出最佳参数和得分 print(f"最佳得分: {grid_result.best_score_}") print(f"最佳参数: {grid_result.best_params_}")
2. 使用Keras Tuner(推荐,对多输入模型更友好)
Keras Tuner是TensorFlow官方推出的超参数调优工具,完美支持Functional API的多输入模型,比GridSearchCV更高效(支持随机搜索、贝叶斯优化等)。
安装Keras Tuner
pip install keras-tuner
定义HyperModel并搜索
import keras_tuner as kt class MultiInputHyperModel(kt.HyperModel): def __init__(self, input_shape1, input_shape2): self.input_shape1 = input_shape1 self.input_shape2 = input_shape2 def build(self, hp): # 定义输入层 input1 = Input(shape=self.input_shape1, name="Input1") input2 = Input(shape=self.input_shape2, name="Input2") # 第一个分支的超参数搜索 units1 = hp.Int('units1', min_value=5, max_value=32, step=5) x = BatchNormalization()(input1) x = Dense(units=units1, activation='relu', kernel_regularizer=l2(0.001))(x) # 第二个分支的超参数搜索 units2 = hp.Int('units2', min_value=32, max_value=128, step=32) y = BatchNormalization()(input2) y = Dense(units=units2, activation="relu", kernel_regularizer=l2(0.001))(y) # 合并分支 concatenated = Concatenate()([x, y]) # Dropout超参数 dropout_rate = hp.Float('dropout_rate', min_value=0.05, max_value=0.2, step=0.03) outputs = Dropout(dropout_rate)(concatenated) outputs = Dense(1, name="output")(outputs) # 优化器超参数 lr = hp.Float('lr', min_value=1e-4, max_value=1e-2, sampling='log') final_lr = hp.Float('final_lr', min_value=0.05, max_value=0.2, step=0.05) model = Model(inputs=[input1, input2], outputs=outputs) model.compile(loss='mse', optimizer=AdaBound(lr=lr, final_lr=final_lr), metrics=['mse']) return model # 初始化超模型 hypermodel = MultiInputHyperModel(input_shape1=X1.shape[1:], input_shape2=X2.shape[1:]) # 初始化调优器(这里用贝叶斯优化) tuner = kt.BayesianOptimization( hypermodel, objective='val_mse', max_trials=20, executions_per_trial=2, directory='tuner_dir', project_name='multi_input_regression' ) # 运行搜索 tuner.search([X1, X2], y, epochs=50, validation_split=0.2, verbose=1) # 获取最佳模型 best_model = tuner.get_best_models(num_models=1)[0] best_hps = tuner.get_best_hyperparameters(num_trials=1)[0] # 打印最佳超参数 print(f"最佳units1: {best_hps.get('units1')}") print(f"最佳units2: {best_hps.get('units2')}") print(f"最佳dropout_rate: {best_hps.get('dropout_rate')}") print(f"最佳lr: {best_hps.get('lr')}") print(f"最佳final_lr: {best_hps.get('final_lr')}")
3. 直接用Hyperopt替代Hyperas
Hyperas本质是Hyperopt的封装,如果你觉得Hyperas不好用,可以直接用Hyperopt来定义超参数空间,手动构建模型并评估,灵活性更高。核心思路是定义一个目标函数,在函数里根据Hyperopt给出的超参数构建模型、训练、返回损失,然后用fmin进行搜索。
关键注意点
- 原来的
build_model函数不要把输入数据作为参数传入,模型构建应该只依赖输入形状和超参数,数据应该在训练时传入fit方法。 - 多输入模型在
fit时要把输入数据以列表的形式传入,比如model.fit([X1, X2], y, ...),这一点在所有调参方法里都要注意。
内容的提问来源于stack exchange,提问作者machine_apprentice
相关产品推荐
相关产品推荐

