You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

K折训练后GPU内存不足,如何有效清理GPU内存?

GPU内存溢出问题解决(10折交叉验证场景)

问题描述

10折交叉验证训练过程中,完成若干折后触发GPU内存不足错误,报错信息如下:

Allocator (GPU_0_bfc) ran out of memory trying to allocate 16.15GiB (rounded to 17338417152)requested by op _EagerConst

系统提示若为内存碎片导致,可设置环境变量TF_GPU_ALLOCATOR=cuda_malloc_async优化。通过nvidia-smi观察到:第一块GPU占用33/48GB,其余3块仅占用13/48GB。已尝试tf.keras.backend.clear_session()但未解决问题,训练代码如下:

physical_devices = tf.config.list_physical_devices("GPU")
for device in physical_devices:
    tf.config.experimental.set_memory_growth(device, True)

strategy = tf.distribute.MirroredStrategy()
print('Number of devices: {}'.format(strategy.num_replicas_in_sync))
..
model code
..
#k fold training loop 
num_folds = 10
kfold = KFold(n_splits=num_folds, shuffle=True)
fold = 0
for train, val in kfold.split(X_train, y_train_one_hot):
    with strategy.scope():
        model = create_model()

    checkpoint = ModelCheckpoint(f'classifier_with_attention_model.{fold + 1}.h5', save_best_only=True)
    history = model.fit(X_train[train], y_train_one_hot[train], epochs=100, batch_size=64,
                        validation_data=(X_train[val], y_train_one_hot[val]), callbacks=[checkpoint])
    tf.keras.backend.clear_session()
    scores = model.evaluate(X_train[val], y_train_one_hot[val])
    fold += 1
    print(f"Fold: {fold}, Loss: {scores[0]}, Accuracy: {scores[1]}")
    cvscores.append(scores[1] * 100)
    history_list.append(history)

解决方案

1. 配置GPU内存分配器环境变量

按系统提示设置环境变量,解决内存碎片问题:

  • Linux/macOS终端:
    export TF_GPU_ALLOCATOR=cuda_malloc_async
    python your_training_script.py
    
  • Windows命令提示符:
    set TF_GPU_ALLOCATOR=cuda_malloc_async
    python your_training_script.py
    
  • Windows PowerShell:
    $env:TF_GPU_ALLOCATOR="cuda_malloc_async"
    python your_training_script.py
    

2. 彻底释放循环内的内存引用

当前代码中model、history等对象未被显式删除,导致clear_session()无法彻底释放显存,调整循环逻辑:

num_folds = 10
kfold = KFold(n_splits=num_folds, shuffle=True)
fold = 0
cvscores = []
history_list = []

for train, val in kfold.split(X_train, y_train_one_hot):
    with strategy.scope():
        model = create_model()

    checkpoint = ModelCheckpoint(f'classifier_with_attention_model.{fold + 1}.h5', save_best_only=True)
    history = model.fit(X_train[train], y_train_one_hot[train], epochs=100, batch_size=64,
                        validation_data=(X_train[val], y_train_one_hot[val]), callbacks=[checkpoint])
    
    # 先完成评估再执行内存释放
    scores = model.evaluate(X_train[val], y_train_one_hot[val])
    fold += 1
    print(f"Fold: {fold}, Loss: {scores[0]}, Accuracy: {scores[1]}")
    cvscores.append(scores[1] * 100)
    history_list.append(history)
    
    # 显式删除对象+强制垃圾回收
    del model
    del history
    tf.keras.backend.clear_session()
    import gc
    gc.collect()

3. 优化批量大小与数据加载

  • 降低单GPU批量:使用MirroredStrategy时,实际批量为batch_size * 设备数,当前batch_size=64对应4块GPU总批量256,可尝试改为32或16。
  • 改用tf.data.Dataset加载数据,避免全量数据驻留内存:
    def create_dataset(X, y, batch_size=32):
        dataset = tf.data.Dataset.from_tensor_slices((X, y))
        dataset = dataset.shuffle(buffer_size=len(X)).batch(batch_size).prefetch(tf.data.AUTOTUNE)
        return dataset
    
    # 循环内替换数据加载方式
    train_dataset = create_dataset(X_train[train], y_train_one_hot[train], batch_size=32)
    val_dataset = create_dataset(X_train[val], y_train_one_hot[val], batch_size=32)
    history = model.fit(train_dataset, epochs=100, validation_data=val_dataset, callbacks=[checkpoint])
    

4. 启用混合精度训练

通过混合精度减少显存占用,需确保输出层使用float32避免精度损失:

from tensorflow.keras import mixed_precision
mixed_precision.set_global_policy('mixed_float16')

# 在create_model()的输出层添加dtype参数
output_layer = tf.keras.layers.Dense(num_classes, activation='softmax', dtype='float32')

5. 均衡多GPU负载

GPU0占用过高可能是数据或模型变量分布不均,可尝试:

  • 确认MirroredStrategy初始化时未限制设备范围,确保所有GPU参与变量同步。
  • 检查数据切片逻辑,保证训练/验证数据在各GPU间均匀分配。

内容的提问来源于stack exchange,提问作者Subhram Dasgupta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 13:20:03