K折训练后GPU内存不足,如何有效清理GPU内存?
GPU内存溢出问题解决(10折交叉验证场景)
问题描述
10折交叉验证训练过程中,完成若干折后触发GPU内存不足错误,报错信息如下:
Allocator (GPU_0_bfc) ran out of memory trying to allocate 16.15GiB (rounded to 17338417152)requested by op _EagerConst
系统提示若为内存碎片导致,可设置环境变量TF_GPU_ALLOCATOR=cuda_malloc_async优化。通过nvidia-smi观察到:第一块GPU占用33/48GB,其余3块仅占用13/48GB。已尝试tf.keras.backend.clear_session()但未解决问题,训练代码如下:
physical_devices = tf.config.list_physical_devices("GPU") for device in physical_devices: tf.config.experimental.set_memory_growth(device, True) strategy = tf.distribute.MirroredStrategy() print('Number of devices: {}'.format(strategy.num_replicas_in_sync)) .. model code .. #k fold training loop num_folds = 10 kfold = KFold(n_splits=num_folds, shuffle=True) fold = 0 for train, val in kfold.split(X_train, y_train_one_hot): with strategy.scope(): model = create_model() checkpoint = ModelCheckpoint(f'classifier_with_attention_model.{fold + 1}.h5', save_best_only=True) history = model.fit(X_train[train], y_train_one_hot[train], epochs=100, batch_size=64, validation_data=(X_train[val], y_train_one_hot[val]), callbacks=[checkpoint]) tf.keras.backend.clear_session() scores = model.evaluate(X_train[val], y_train_one_hot[val]) fold += 1 print(f"Fold: {fold}, Loss: {scores[0]}, Accuracy: {scores[1]}") cvscores.append(scores[1] * 100) history_list.append(history)
解决方案
1. 配置GPU内存分配器环境变量
按系统提示设置环境变量,解决内存碎片问题:
- Linux/macOS终端:
export TF_GPU_ALLOCATOR=cuda_malloc_async python your_training_script.py - Windows命令提示符:
set TF_GPU_ALLOCATOR=cuda_malloc_async python your_training_script.py - Windows PowerShell:
$env:TF_GPU_ALLOCATOR="cuda_malloc_async" python your_training_script.py
2. 彻底释放循环内的内存引用
当前代码中model、history等对象未被显式删除,导致clear_session()无法彻底释放显存,调整循环逻辑:
num_folds = 10 kfold = KFold(n_splits=num_folds, shuffle=True) fold = 0 cvscores = [] history_list = [] for train, val in kfold.split(X_train, y_train_one_hot): with strategy.scope(): model = create_model() checkpoint = ModelCheckpoint(f'classifier_with_attention_model.{fold + 1}.h5', save_best_only=True) history = model.fit(X_train[train], y_train_one_hot[train], epochs=100, batch_size=64, validation_data=(X_train[val], y_train_one_hot[val]), callbacks=[checkpoint]) # 先完成评估再执行内存释放 scores = model.evaluate(X_train[val], y_train_one_hot[val]) fold += 1 print(f"Fold: {fold}, Loss: {scores[0]}, Accuracy: {scores[1]}") cvscores.append(scores[1] * 100) history_list.append(history) # 显式删除对象+强制垃圾回收 del model del history tf.keras.backend.clear_session() import gc gc.collect()
3. 优化批量大小与数据加载
- 降低单GPU批量:使用
MirroredStrategy时,实际批量为batch_size * 设备数,当前batch_size=64对应4块GPU总批量256,可尝试改为32或16。 - 改用
tf.data.Dataset加载数据,避免全量数据驻留内存:def create_dataset(X, y, batch_size=32): dataset = tf.data.Dataset.from_tensor_slices((X, y)) dataset = dataset.shuffle(buffer_size=len(X)).batch(batch_size).prefetch(tf.data.AUTOTUNE) return dataset # 循环内替换数据加载方式 train_dataset = create_dataset(X_train[train], y_train_one_hot[train], batch_size=32) val_dataset = create_dataset(X_train[val], y_train_one_hot[val], batch_size=32) history = model.fit(train_dataset, epochs=100, validation_data=val_dataset, callbacks=[checkpoint])
4. 启用混合精度训练
通过混合精度减少显存占用,需确保输出层使用float32避免精度损失:
from tensorflow.keras import mixed_precision mixed_precision.set_global_policy('mixed_float16') # 在create_model()的输出层添加dtype参数 output_layer = tf.keras.layers.Dense(num_classes, activation='softmax', dtype='float32')
5. 均衡多GPU负载
GPU0占用过高可能是数据或模型变量分布不均,可尝试:
- 确认
MirroredStrategy初始化时未限制设备范围,确保所有GPU参与变量同步。 - 检查数据切片逻辑,保证训练/验证数据在各GPU间均匀分配。
内容的提问来源于stack exchange,提问作者Subhram Dasgupta
相关产品推荐
相关产品推荐

