You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Keras Functional API的多输入回归模型超参数调优求助

解决Keras Functional多输入模型的超参数调优问题

我明白你现在的困境——用Functional API搭建了多输入回归模型,但GridSearchCV和Hyperas没法直接支持,单独调分支又不合理。其实有几个实用的方案能搞定这个问题,下面给你详细说:

1. 用KerasRegressor包装模型适配GridSearchCV

GridSearchCV需要符合scikit-learn的estimator接口,所以我们可以把Functional模型用KerasRegressor包装起来,关键是要把超参数作为参数传入模型构建函数,而不是像你原来那样把输入数据传进去(原来的build_model设计有问题,模型构建不该依赖具体数据)。

修改模型构建函数

先重构你的build_model,让它接受超参数,输入层的shape从外部传入:

from tensorflow.keras.models import Model
from tensorflow.keras.layers import Input, Dense, BatchNormalization, Concatenate, Dropout
from tensorflow.keras.regularizers import l2
from tensorflow.keras.optimizers import AdaBound

def build_model(input_shape1, input_shape2, units1=5, units2=64, dropout_rate=0.08, lr=0.001, final_lr=0.1):
    """
    创建多输入ANN模型,接受超参数和输入形状
    :param input_shape1: 第一个输入的形状(不含batch维度)
    :param input_shape2: 第二个输入的形状(不含batch维度)
    :param units1: 第一个分支Dense层的单元数
    :param units2: 第二个分支Dense层的单元数
    :param dropout_rate: Dropout层的丢弃率
    :param lr: 优化器初始学习率
    :param final_lr: AdaBound的最终学习率
    :return: 编译好的模型
    """
    # 定义输入层
    input1 = Input(shape=input_shape1, name="Input1")
    input2 = Input(shape=input_shape2, name="Input2")
    
    # 第一个分支
    x = BatchNormalization()(input1)
    x = Dense(units=units1, activation='relu', kernel_regularizer=l2(0.001))(x)
    
    # 第二个分支
    y = BatchNormalization()(input2)
    y = Dense(units=units2, activation="relu", kernel_regularizer=l2(0.001))(y)
    
    # 合并分支
    concatenated = Concatenate()([x, y])
    
    # 输出层
    outputs = Dropout(dropout_rate)(concatenated)
    outputs = Dense(1, name="output")(outputs)
    
    # 创建并编译模型
    model = Model(inputs=[input1, input2], outputs=outputs)
    model.compile(loss='mse', optimizer=AdaBound(lr=lr, final_lr=final_lr), metrics=['mse'])
    
    return model

用KerasRegressor包装并做GridSearchCV

然后定义一个包装函数,把超参数传递给build_model,再用GridSearchCV搜索:

from tensorflow.keras.wrappers.scikit_learn import KerasRegressor
from sklearn.model_selection import GridSearchCV
import numpy as np

# 假设你的输入数据是X1, X2,标签是y
# 先处理数据形状(这里假设你已经做好了预处理)
input_shape1 = X1.shape[1:]
input_shape2 = X2.shape[1:]

# 定义包装函数
def create_model(input_shape1, input_shape2):
    def wrapper(units1, units2, dropout_rate, lr, final_lr):
        return build_model(input_shape1, input_shape2, units1=units1, units2=units2, dropout_rate=dropout_rate, lr=lr, final_lr=final_lr)
    return wrapper

# 初始化KerasRegressor
model = KerasRegressor(build_fn=create_model(input_shape1, input_shape2), verbose=1)

# 设置超参数搜索空间
param_grid = {
    'units1': [5, 10, 16],
    'units2': [32, 64, 128],
    'dropout_rate': [0.05, 0.08, 0.1],
    'lr': [0.0001, 0.001, 0.01],
    'final_lr': [0.05, 0.1, 0.2]
}

# 运行GridSearchCV
grid = GridSearchCV(estimator=model, param_grid=param_grid, cv=3, scoring='neg_mean_squared_error')
grid_result = grid.fit([X1, X2], y)

# 输出最佳参数和得分
print(f"最佳得分: {grid_result.best_score_}")
print(f"最佳参数: {grid_result.best_params_}")

2. 使用Keras Tuner(推荐,对多输入模型更友好)

Keras Tuner是TensorFlow官方推出的超参数调优工具,完美支持Functional API的多输入模型,比GridSearchCV更高效(支持随机搜索、贝叶斯优化等)。

安装Keras Tuner

pip install keras-tuner

定义HyperModel并搜索

import keras_tuner as kt

class MultiInputHyperModel(kt.HyperModel):
    def __init__(self, input_shape1, input_shape2):
        self.input_shape1 = input_shape1
        self.input_shape2 = input_shape2
    
    def build(self, hp):
        # 定义输入层
        input1 = Input(shape=self.input_shape1, name="Input1")
        input2 = Input(shape=self.input_shape2, name="Input2")
        
        # 第一个分支的超参数搜索
        units1 = hp.Int('units1', min_value=5, max_value=32, step=5)
        x = BatchNormalization()(input1)
        x = Dense(units=units1, activation='relu', kernel_regularizer=l2(0.001))(x)
        
        # 第二个分支的超参数搜索
        units2 = hp.Int('units2', min_value=32, max_value=128, step=32)
        y = BatchNormalization()(input2)
        y = Dense(units=units2, activation="relu", kernel_regularizer=l2(0.001))(y)
        
        # 合并分支
        concatenated = Concatenate()([x, y])
        
        # Dropout超参数
        dropout_rate = hp.Float('dropout_rate', min_value=0.05, max_value=0.2, step=0.03)
        outputs = Dropout(dropout_rate)(concatenated)
        outputs = Dense(1, name="output")(outputs)
        
        # 优化器超参数
        lr = hp.Float('lr', min_value=1e-4, max_value=1e-2, sampling='log')
        final_lr = hp.Float('final_lr', min_value=0.05, max_value=0.2, step=0.05)
        
        model = Model(inputs=[input1, input2], outputs=outputs)
        model.compile(loss='mse', optimizer=AdaBound(lr=lr, final_lr=final_lr), metrics=['mse'])
        
        return model

# 初始化超模型
hypermodel = MultiInputHyperModel(input_shape1=X1.shape[1:], input_shape2=X2.shape[1:])

# 初始化调优器(这里用贝叶斯优化)
tuner = kt.BayesianOptimization(
    hypermodel,
    objective='val_mse',
    max_trials=20,
    executions_per_trial=2,
    directory='tuner_dir',
    project_name='multi_input_regression'
)

# 运行搜索
tuner.search([X1, X2], y, epochs=50, validation_split=0.2, verbose=1)

# 获取最佳模型
best_model = tuner.get_best_models(num_models=1)[0]
best_hps = tuner.get_best_hyperparameters(num_trials=1)[0]

# 打印最佳超参数
print(f"最佳units1: {best_hps.get('units1')}")
print(f"最佳units2: {best_hps.get('units2')}")
print(f"最佳dropout_rate: {best_hps.get('dropout_rate')}")
print(f"最佳lr: {best_hps.get('lr')}")
print(f"最佳final_lr: {best_hps.get('final_lr')}")

3. 直接用Hyperopt替代Hyperas

Hyperas本质是Hyperopt的封装,如果你觉得Hyperas不好用,可以直接用Hyperopt来定义超参数空间,手动构建模型并评估,灵活性更高。核心思路是定义一个目标函数,在函数里根据Hyperopt给出的超参数构建模型、训练、返回损失,然后用fmin进行搜索。

关键注意点

  • 原来的build_model函数不要把输入数据作为参数传入,模型构建应该只依赖输入形状和超参数,数据应该在训练时传入fit方法。
  • 多输入模型在fit时要把输入数据以列表的形式传入,比如model.fit([X1, X2], y, ...),这一点在所有调参方法里都要注意。

内容的提问来源于stack exchange,提问作者machine_apprentice

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 18:32:55