You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Keras Tuner中正确实现batch_size超参数调优?

问题

我是Keras新手,使用KerasTuner进行超参数调优时效果良好,但始终无法成功调优batch_size参数。参考官方讨论后尝试了如下代码,但发现batch_size并未发生变化。请问该功能是否仅支持RandomSearch调优器?自定义fit方法应如何正确配置才能生效?

class ANN: 

 def build_ann(self, hp):

        model = Sequential()

        for i in range(hp.Int('num_layers', 1, 10)): 
            model.add(LSTM(units=hp.Int('units_' + str(i), min_value=2, max_value=20, step=1), return_sequences=(i < hp.Int('num_layers', 1, 10) - 1)))
            model.add(Dropout(rate=hp.Float('dropout_' + str(i), 0, 0.5, step=0.1)))    
        model.add(Dense(units=hp.Int('units_last', min_value=2, max_value=20, step=1)))
        model.add(Dense(units=len(self.targets), activation='sigmoid'))

        model.compile(optimizer=Adam(hp.Choice('learning_rate', values=[1e-2, 1e-3, 1e-4])),
                    loss='mean_squared_error', metrics=['accuracy'])

        model.build(input_shape=(1, self.maxlen, len(self.features)))

        return model
    
    def fit(self, hp, model, *args, **kwargs):
        return model.fit(
            *args,
            batch_size=hp.Int('batch_size', 1, 10, step=16),
            **kwargs,
        )
tuner = keras_tuner.Hyperband(
    ann.build_ann,
    objective='val_accuracy',
    max_epochs=50,
    factor=2,
    overwrite=True,
    directory='my_dir2',
    project_name='my_project')

tuner.search(trainx, trainy, epochs=50, validation_split=0.2 )
回答

核心问题分析

你的代码有两个关键问题导致batch_size不生效:

  1. 自定义fit方法未被识别:你的ANN类没有继承keras_tuner.HyperModel,Keras Tuner的所有调优器(包括Hyperband)都不会自动调用类中的fit方法,只会使用默认的模型训练逻辑。
  2. batch_size的超参数配置错误:hp.Int('batch_size', 1, 10, step=16)中,step值远大于取值范围(1-10),导致只会生成1这一个固定值,自然看不到batch_size变化。

自定义fit方法的正确配置

所有Keras Tuner调优器(Hyperband、RandomSearch、BayesianOptimization等)都支持自定义训练逻辑,不需要局限于RandomSearch。正确做法是让你的类继承keras_tuner.HyperModel,并实现build和fit方法:

import keras_tuner as kt
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import LSTM, Dropout, Dense
from tensorflow.keras.optimizers import Adam

class ANN(kt.HyperModel):
    def __init__(self, maxlen, features, targets):
        self.maxlen = maxlen
        self.features = features
        self.targets = targets

    def build(self, hp):
        model = Sequential()
        num_layers = hp.Int('num_layers', 1, 10)
        for i in range(num_layers): 
            # 避免重复调用hp.Int('num_layers'),先存变量复用
            model.add(LSTM(
                units=hp.Int(f'units_{i}', min_value=2, max_value=20, step=1),
                return_sequences=(i < num_layers - 1)
            ))
            model.add(Dropout(rate=hp.Float(f'dropout_{i}', 0, 0.5, step=0.1)))    
        model.add(Dense(units=hp.Int('units_last', min_value=2, max_value=20, step=1)))
        model.add(Dense(units=len(self.targets), activation='sigmoid'))

        model.compile(
            optimizer=Adam(hp.Choice('learning_rate', values=[1e-2, 1e-3, 1e-4])),
            loss='mean_squared_error', 
            metrics=['accuracy']
        )
        return model
    
    def fit(self, hp, model, *args, **kwargs):
        # 修正step值,比如设置为2,在1-32范围内生成可选batch_size
        return model.fit(
            *args,
            batch_size=hp.Int('batch_size', min_value=1, max_value=32, step=2),
            **kwargs,
        )

然后创建调优器时,传入HyperModel实例,而非单独的build方法:

# 初始化你的HyperModel实例
ann_hypermodel = ANN(maxlen=your_maxlen, features=your_features, targets=your_targets)

tuner = kt.Hyperband(
    ann_hypermodel,  # 传入HyperModel实例
    objective='val_accuracy',
    max_epochs=50,
    factor=2,
    overwrite=True,
    directory='my_dir2',
    project_name='my_project')

tuner.search(trainx, trainy, validation_split=0.2 )

额外优化点

  • 避免在循环中重复调用hp.Int('num_layers'),先将值存入变量复用,减少冗余计算。
  • 根据你的数据集大小调整batch_size的取值范围和step,比如如果数据集较大,可将max_value设为64或128,step设为8,更符合实际训练场景。

内容的提问来源于stack exchange,提问作者AndyEverything

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 12:10:39