You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow 10折交叉验证训练GPU内存溢出问题求助

解决TensorFlow 10折交叉验证GPU内存溢出问题

问题根源分析

你的代码存在几个关键问题导致GPU内存泄漏:

  1. 循环内重复创建模型但未清理旧模型占用的GPU资源
  2. 每次训练/评估时重复将数据集转换为numpy数组,造成内存冗余
  3. fold_no的初始化和更新逻辑错误(不影响内存,但会导致折数显示异常)

具体解决步骤

  • 提前转换数据集为numpy数组:避免循环内重复转换,减少内存分配
  • 循环末尾强制清理模型与GPU内存:释放当前折模型占用的所有资源
  • 开启GPU内存按需分配:让TensorFlow仅使用所需内存,而非一次性占满
  • 修正fold_no的逻辑错误:确保折数正确递增

修正后的完整代码

import numpy as np
import tensorflow as tf
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Conv2D, MaxPooling2D, Flatten, Dense, Dropout
import gc  # 导入垃圾回收模块

# 开启GPU内存按需分配
gpus = tf.config.experimental.list_physical_devices('GPU')
if gpus:
    try:
        for gpu in gpus:
            tf.config.experimental.set_memory_growth(gpu, True)
    except RuntimeError as e:
        print(e)

acc_per_fold = []
loss_per_fold = []
fold_no = 1  # 将fold_no初始化移到循环外

# 提前转换为numpy数组,避免循环内重复操作
x_train_np = np.array(x_train)
y_train_np = np.array(y_train)

for train, test in kfold.split(x_train_np, y_train_np):
    # Define the model architecture
    model = Sequential()
    model.add(Conv2D(32, kernel_size=(3,3), input_shape = x_train_np[0].shape, activation = "relu"))
    model.add(MaxPooling2D(2,2))
    model.add(Conv2D(32, kernel_size=(3,3), activation = "relu"))
    model.add(MaxPooling2D(2,2))

    model.add(Flatten())
    model.add(Dense(64, activation = "relu"))
    model.add(Dropout(0.1))
    model.add(Dense(32, activation = "tanh"))
    model.add(Dense(1, activation = "sigmoid"))

    # Compile the model
    model.compile(loss = "binary_crossentropy", 
              optimizer = tf.keras.optimizers.Adam(learning_rate = 0.001), 
              metrics = ["accuracy"])


    # Generate a print
    print('------------------------------------------------------------------------')
    print(f'Training for fold {fold_no} ...')
    # Fit data to model
    history = model.fit(x_train_np[train], y_train_np[train],
              batch_size=32,
              epochs=10,
              verbose=1)

    # Generate generalization metrics
    scores = model.evaluate(x_train_np[test], y_train_np[test], verbose=0)
    print(f"Score for fold {fold_no}: {model.metrics_names[0]} of {scores[0]}; {model.metrics_names[1]} of {scores[1]*100}%")
    acc_per_fold.append(scores[1] * 100)
    loss_per_fold.append(scores[0])

    # 清理当前折的资源
    del model, history
    tf.keras.backend.clear_session()
    gc.collect()

    # 正确递增折数
    fold_no += 1

额外说明

  • tf.keras.backend.clear_session()会清除当前Keras会话的所有模型和层,释放GPU内存
  • gc.collect()强制Python进行垃圾回收,清理未被引用的内存对象
  • 内存按需分配设置只需在代码开头执行一次,无需循环内重复运行

内容的提问来源于stack exchange,提问作者user20133182

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 21:40:38