TensorFlow 10折交叉验证训练GPU内存溢出问题求助
解决TensorFlow 10折交叉验证GPU内存溢出问题
问题根源分析
你的代码存在几个关键问题导致GPU内存泄漏:
- 循环内重复创建模型但未清理旧模型占用的GPU资源
- 每次训练/评估时重复将数据集转换为numpy数组,造成内存冗余
- fold_no的初始化和更新逻辑错误(不影响内存,但会导致折数显示异常)
具体解决步骤
- 提前转换数据集为numpy数组:避免循环内重复转换,减少内存分配
- 循环末尾强制清理模型与GPU内存:释放当前折模型占用的所有资源
- 开启GPU内存按需分配:让TensorFlow仅使用所需内存,而非一次性占满
- 修正fold_no的逻辑错误:确保折数正确递增
修正后的完整代码
import numpy as np import tensorflow as tf from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Conv2D, MaxPooling2D, Flatten, Dense, Dropout import gc # 导入垃圾回收模块 # 开启GPU内存按需分配 gpus = tf.config.experimental.list_physical_devices('GPU') if gpus: try: for gpu in gpus: tf.config.experimental.set_memory_growth(gpu, True) except RuntimeError as e: print(e) acc_per_fold = [] loss_per_fold = [] fold_no = 1 # 将fold_no初始化移到循环外 # 提前转换为numpy数组,避免循环内重复操作 x_train_np = np.array(x_train) y_train_np = np.array(y_train) for train, test in kfold.split(x_train_np, y_train_np): # Define the model architecture model = Sequential() model.add(Conv2D(32, kernel_size=(3,3), input_shape = x_train_np[0].shape, activation = "relu")) model.add(MaxPooling2D(2,2)) model.add(Conv2D(32, kernel_size=(3,3), activation = "relu")) model.add(MaxPooling2D(2,2)) model.add(Flatten()) model.add(Dense(64, activation = "relu")) model.add(Dropout(0.1)) model.add(Dense(32, activation = "tanh")) model.add(Dense(1, activation = "sigmoid")) # Compile the model model.compile(loss = "binary_crossentropy", optimizer = tf.keras.optimizers.Adam(learning_rate = 0.001), metrics = ["accuracy"]) # Generate a print print('------------------------------------------------------------------------') print(f'Training for fold {fold_no} ...') # Fit data to model history = model.fit(x_train_np[train], y_train_np[train], batch_size=32, epochs=10, verbose=1) # Generate generalization metrics scores = model.evaluate(x_train_np[test], y_train_np[test], verbose=0) print(f"Score for fold {fold_no}: {model.metrics_names[0]} of {scores[0]}; {model.metrics_names[1]} of {scores[1]*100}%") acc_per_fold.append(scores[1] * 100) loss_per_fold.append(scores[0]) # 清理当前折的资源 del model, history tf.keras.backend.clear_session() gc.collect() # 正确递增折数 fold_no += 1
额外说明
tf.keras.backend.clear_session()会清除当前Keras会话的所有模型和层,释放GPU内存gc.collect()强制Python进行垃圾回收,清理未被引用的内存对象- 内存按需分配设置只需在代码开头执行一次,无需循环内重复运行
内容的提问来源于stack exchange,提问作者user20133182
相关产品推荐
相关产品推荐

