Keras训练循环GPU清理及多模型训练独立性实现方法
问题描述
我正在探索在单个代码库中训练多个模型的方法,目标是通过不同随机种子生成各类模型,评估多样的网络架构与训练数据。训练数据通过随机种子生成以实现多样的数据划分,但发现单模型训练与多模型训练的结果存在差异,怀疑初始模型训练会影响后续训练流程。请问如何在Keras训练循环中清理GPU,并确保各模型训练过程完全独立?
现有代码片段
def build_model(self, seed): print("Creating Model With Seed:", seed) regularizer = keras.regularizers.L1L2(l1=1e-5, l2=1e-4) initializer = keras.initializers.GlorotUniform(seed=seed) model = keras.Sequential([ layers.Input(shape=self.input_shape), layers.Dense(6, activation='sigmoid', kernel_initializer=initializer, kernel_regularizer=regularizer), layers.Dense(1, activation='linear', kernel_initializer=initializer) ]) model.compile(loss=lambda y_true, y_pred: self.huber_loss(y_true, y_pred, delta=0.35), optimizer='adam', metrics=['mae', 'mse']) return model # 模型训练函数 def train_models(self, model, seed, epochs, batch_size): early_stop = keras.callbacks.EarlyStopping(monitor='val_loss', patience=50) histories = [] print("Training Model With Seed Number:", seed) for i, seed in enumerate(self.seeds): # 训练模型 history = model.fit( self.normed_train_data[i], self.train_labels[i], epochs=epochs, validation_split=0.2, verbose=1, batch_size=batch_size, callbacks=[PrintDot(), early_stop] ) histories.append(history) # 重置TensorFlow随机种子为默认值 # tf.random.set_seed(None) return histories
解决方案
1. 修复训练循环的逻辑错误
当前train_models函数在循环第一次迭代就执行return,导致仅训练第一个种子的模型。需将return移出循环,且每个种子对应全新模型实例,不能复用同一个模型对象:
def train_models(self, epochs, batch_size): early_stop = keras.callbacks.EarlyStopping(monitor='val_loss', patience=50) histories = [] for i, seed in enumerate(self.seeds): print("Training Model With Seed Number:", seed) # 为每个种子构建全新模型 model = self.build_model(seed) # 固定全局随机种子,消除跨训练的随机干扰 self.set_global_seeds(seed) history = model.fit( self.normed_train_data[i], self.train_labels[i], epochs=epochs, validation_split=0.2, verbose=1, batch_size=batch_size, callbacks=[PrintDot(), early_stop] ) histories.append(history) # 清理当前模型的GPU资源 self.cleanup_model(model) return histories
2. 固定全局随机种子
添加全局种子设置函数,确保每个训练过程的随机性完全受控:
def set_global_seeds(self, seed): import random import numpy as np random.seed(seed) np.random.seed(seed) tf.random.set_seed(seed) # 启用确定性操作,避免CuDNN的非随机行为 tf.config.experimental.enable_op_determinism()
3. 显式清理GPU内存
训练完每个模型后,强制释放显存资源,避免前序模型占用影响后续训练:
def cleanup_model(self, model): # 删除模型对象 del model # 清除Keras会话,释放GPU内存 tf.keras.backend.clear_session() # 重置TensorFlow默认图 tf.compat.v1.reset_default_graph() # 触发Python垃圾回收 import gc gc.collect()
4. 验证训练独立性
训练完成后,对比单模型单独训练的指标(损失、MAE、MSE)与多模型循环训练的对应指标,确认数值一致,即可证明训练过程无交叉干扰。
内容的提问来源于stack exchange,提问作者Agenor Maradiaga
相关产品推荐
相关产品推荐

