You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras训练循环GPU清理及多模型训练独立性实现方法

问题描述

我正在探索在单个代码库中训练多个模型的方法,目标是通过不同随机种子生成各类模型,评估多样的网络架构与训练数据。训练数据通过随机种子生成以实现多样的数据划分,但发现单模型训练与多模型训练的结果存在差异,怀疑初始模型训练会影响后续训练流程。请问如何在Keras训练循环中清理GPU,并确保各模型训练过程完全独立?

现有代码片段
def build_model(self, seed):
    print("Creating Model With Seed:", seed)
    regularizer = keras.regularizers.L1L2(l1=1e-5, l2=1e-4)
    initializer = keras.initializers.GlorotUniform(seed=seed)
    model = keras.Sequential([
        layers.Input(shape=self.input_shape),
        layers.Dense(6, activation='sigmoid', kernel_initializer=initializer, kernel_regularizer=regularizer),
        layers.Dense(1, activation='linear', kernel_initializer=initializer)
    ])

    model.compile(loss=lambda y_true, y_pred: self.huber_loss(y_true, y_pred, delta=0.35),
                  optimizer='adam',
                  metrics=['mae', 'mse'])
    return model

# 模型训练函数
def train_models(self, model, seed, epochs, batch_size):
    early_stop = keras.callbacks.EarlyStopping(monitor='val_loss', patience=50)
    histories = []
    print("Training Model With Seed Number:", seed)

    for i, seed in enumerate(self.seeds):
        # 训练模型
        history = model.fit(
            self.normed_train_data[i],
            self.train_labels[i],
            epochs=epochs, validation_split=0.2, verbose=1, batch_size=batch_size,
            callbacks=[PrintDot(), early_stop]
        )
        histories.append(history)
        # 重置TensorFlow随机种子为默认值
        # tf.random.set_seed(None)
        return histories
解决方案

1. 修复训练循环的逻辑错误

当前train_models函数在循环第一次迭代就执行return,导致仅训练第一个种子的模型。需将return移出循环,且每个种子对应全新模型实例,不能复用同一个模型对象:

def train_models(self, epochs, batch_size):
    early_stop = keras.callbacks.EarlyStopping(monitor='val_loss', patience=50)
    histories = []
    for i, seed in enumerate(self.seeds):
        print("Training Model With Seed Number:", seed)
        # 为每个种子构建全新模型
        model = self.build_model(seed)
        # 固定全局随机种子,消除跨训练的随机干扰
        self.set_global_seeds(seed)
        history = model.fit(
            self.normed_train_data[i],
            self.train_labels[i],
            epochs=epochs, validation_split=0.2, verbose=1, batch_size=batch_size,
            callbacks=[PrintDot(), early_stop]
        )
        histories.append(history)
        # 清理当前模型的GPU资源
        self.cleanup_model(model)
    return histories

2. 固定全局随机种子

添加全局种子设置函数,确保每个训练过程的随机性完全受控:

def set_global_seeds(self, seed):
    import random
    import numpy as np
    random.seed(seed)
    np.random.seed(seed)
    tf.random.set_seed(seed)
    # 启用确定性操作,避免CuDNN的非随机行为
    tf.config.experimental.enable_op_determinism()

3. 显式清理GPU内存

训练完每个模型后,强制释放显存资源,避免前序模型占用影响后续训练:

def cleanup_model(self, model):
    # 删除模型对象
    del model
    # 清除Keras会话,释放GPU内存
    tf.keras.backend.clear_session()
    # 重置TensorFlow默认图
    tf.compat.v1.reset_default_graph()
    # 触发Python垃圾回收
    import gc
    gc.collect()

4. 验证训练独立性

训练完成后,对比单模型单独训练的指标(损失、MAE、MSE)与多模型循环训练的对应指标,确认数值一致,即可证明训练过程无交叉干扰。


内容的提问来源于stack exchange,提问作者Agenor Maradiaga

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 16:34:50