You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于Keras Tuner调优MLP超参数的技术疑问

Keras Tuner超参数调优常见问题解答

调优代码

def build_model2(hp):
      model = tf.keras.Sequential()
      for i in range(hp.Int('layers', 2, 6)):
          model.add(tf.keras.layers.Dense(units=hp.Int('units_' + str(i), 32, 512, step=128), 
                                          activation = hp.Choice('act_' + str(i), ['relu', 'sigmoid','tanh'])))
      model.add(tf.keras.layers.Flatten())
      model.add(tf.keras.layers.Dense(5, activation='softmax'))
      learning_rate = hp.Float("lr", min_value=1e-4, max_value=1e-2, sampling="log")
      model.compile(tf.keras.optimizers.Adam(learning_rate=learning_rate), loss = 'categorical_crossentropy', metrics = ['accuracy'])
      return model
 
tuner2 = tf.keras.tuners.RandomSearch(build_model2, objective = 'val_accuracy', max_trials = 5,
                      executions_per_trial = 3, overwrite=True)
  
tuner2.search_space_summary()
  
tuner2.search(X_train, Y_train, epochs=25, validation_data=(X_train, Y_train),verbose = 1)

tuner2.results_summary()

# Get the optimal hyperparameters
best_hps=tuner2.get_best_hyperparameters(num_trials=1)[0]
print("The optimal parameters are:")
print(best_hps.values)  

# Build the model with the optimal hyperparameters and train it on the data for 50 epochs
model = tuner2.hypermodel.build(best_hps)
history = model.fit(X_train, Y_train, epochs=50, validation_split=0.2)

val_acc_per_epoch = history.history['val_accuracy']
best_epoch = val_acc_per_epoch.index(max(val_acc_per_epoch)) + 1
print('Best epoch: %d' % (best_epoch,))
  
hypermodel = tuner2.hypermodel.build(best_hps)

# Retrain the model
hypermodel.fit(X_train, Y_train, epochs=best_epoch)
  
eval_result = hypermodel.evaluate(X_test, Y_test)
print("[test loss, test accuracy]:", eval_result)

本次调优的参数范围

  • 隐藏层数:2-6层
  • 隐藏层神经元数:32到512,步长128
  • 激活函数:可选relu、sigmoid、tanh
  • 学习率:1e-4到1e-2,采用对数采样

问题1:如何判定某一超参数组合为最优?

Keras Tuner完全依据初始化调优器时指定的目标指标(这里是val_accuracy)判定最优。你的代码里RandomSearch设置了objective='val_accuracy',每个超参数组合会执行3次训练(executions_per_trial=3),取这3次的验证准确率平均值作为该组合的最终得分。所有组合按得分从高到低排序,得分最高的那组就是最优超参数组合。

问题2:若最优隐藏层数为2,为何仍会出现unit_2、unit_3、unit_4的参数值?

这是因为Keras Tuner在构建搜索空间时,会预先注册所有可能的超参数——包括对应3-6层的unit_2、unit_3、unit_4等。即使最终最优组合只用到2层,这些参数依然会被记录在超参数对象中,但实际构建模型时,循环只会执行2次(range(2)),后面的unit_2等参数根本不会被调用,属于无效的遗留项,不会影响模型的结构和性能,完全可以忽略。

问题3:重复模拟后超参数值发生变化,该如何确定最优组合?

这种情况是正常的,主要源于两方面随机性:一是随机搜索本身的采样随机性,二是模型训练过程中的权重初始化、数据shuffle等随机性。可以通过以下方法确定稳定的最优组合:

  • 增加max_trials值,让搜索覆盖更多参数组合,降低随机采样的偏差
  • 提高executions_per_trial值,增加每个参数组合的训练次数,用平均指标来稳定评估结果
  • 多次调优后,优先选择在多次实验中稳定出现且验证/测试指标表现优异的参数组合;也可以把多次调优得到的最优参数对应的模型都训练一遍,最终选取测试集表现最好的那个模型
  • 条件允许的话,改用BayesianOptimization替代随机搜索,它会基于之前的实验结果智能采样,搜索效率更高,结果也更稳定

内容的提问来源于stack exchange,提问作者A. Gehani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 14:54:57