You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow识别GPU但训练时内存暴涨崩溃问题求助

TensorFlow训练时GPU内存暴涨崩溃问题解决

问题现象

  • 系统可正常识别NVIDIA GeForce GTX 960 4GB GPU,执行tf.config.list_logical_devices('GPU')能输出正常设备信息
  • 运行《AI and Machine Learning for Coders》示例代码时,调用model.fit()方法卡在epoch 1/50,内存占用暴涨后程序崩溃
  • 报错信息:Process finished with exit code -1073740791 (0xC0000409),同时出现提示:I tensorflow/stream_executor/cuda/cuda_dnn.cc:384] Loaded cuDNN version 8401

环境配置

  • tensorflow==2.9.0
  • tensorflow-gpu==2.1.0
  • CUDA v11.7
  • CUDNN v8.4.1.50

测试代码

import tensorflow as tf
data = tf.keras.datasets.fashion_mnist

class myCallback(tf.keras.callbacks.Callback):
  def on_epoch_end(self, epoch, logs={}):
    if(logs.get('accuracy')>0.99):
      print("\nReached 99% accuracy so cancelling training!")
      self.model.stop_training = True


callbacks = myCallback()


(training_images, training_labels), (test_images, test_labels) = data.load_data()
print(type(training_images))
training_images=training_images.reshape(60000, 28, 28, 1)
training_images  = training_images / 255.0
test_images = test_images.reshape(10000, 28, 28, 1)
test_images = test_images / 255.0

model = tf.keras.models.Sequential([
    tf.keras.layers.Conv2D(64, (3, 3), activation='relu', input_shape=(28, 28, 1)),
    tf.keras.layers.MaxPooling2D(2, 2),
    tf.keras.layers.Conv2D(64, (3, 3), activation='relu'),
    tf.keras.layers.MaxPooling2D(2,2),
    tf.keras.layers.Flatten(),
    tf.keras.layers.Dense(128, activation=tf.nn.relu),
    tf.keras.layers.Dense(10, activation=tf.nn.softmax)])

model.compile(optimizer='adam',
              loss='sparse_categorical_crossentropy',
              metrics=['accuracy'])

model.summary()

model.fit(training_images, training_labels, epochs=50, callbacks=[callbacks], verbose=1)

tf.config.list_logical_devices('GPU')输出

2022-09-10 23:30:04.052968: I tensorflow/core/platform/cpu_feature_guard.cc:193] This TensorFlow binary is optimized with oneAPI Deep Neural Network Library (oneDNN) to use the following CPU instructions in performance-critical operations:  AVX AVX2
To enable them in other operations, rebuild TensorFlow with the appropriate compiler flags.


2022-09-10 23:30:04.496787: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1616] Created device /job:localhost/replica:0/task:0/device:GPU:0 with 2810 MB memory:  -> device: 0, name: NVIDIA GeForce GTX 960, pci bus id: 0000:06:00.0, compute capability: 5.2
[LogicalDevice(name='/device:GPU:0', device_type='GPU')]

解决方案

1. 清理冲突的TensorFlow版本

同时安装tensorflow==2.9.0和tensorflow-gpu==2.1.0会引发版本冲突,这是核心问题之一:

  • 卸载现有TensorFlow包:pip uninstall tensorflow tensorflow-gpu -y
  • 重新安装单一版本:TensorFlow 2.4+之后,tensorflow包已包含GPU支持,直接安装匹配版本:pip install tensorflow==2.9.0

2. 限制GPU内存按需分配

GTX 960仅4GB显存,默认TensorFlow会占用全部GPU内存,导致溢出。在导入tensorflow后立即添加以下代码:

gpus = tf.config.experimental.list_physical_devices('GPU')
if gpus:
  try:
    for gpu in gpus:
      tf.config.experimental.set_memory_growth(gpu, True)
    logical_gpus = tf.config.experimental.list_logical_devices('GPU')
    print(len(gpus), "Physical GPUs,", len(logical_gpus), "Logical GPUs")
  except RuntimeError as e:
    print(e)

3. 调整训练批次大小

默认model.fit()的batch size为32,可适当减小以降低显存占用:

model.fit(training_images, training_labels, epochs=50, callbacks=[callbacks], verbose=1, batch_size=16)

4. 验证CUDA与TensorFlow兼容性

TensorFlow 2.9.0官方推荐搭配CUDA 11.2、cuDNN 8.1.0,当前CUDA 11.7虽可兼容,但如果问题仍存在,可降级CUDA到11.2版本,并确保cuDNN版本对应匹配。

内容的提问来源于stack exchange,提问作者9879ypxkj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 15:40:23