TensorFlow识别GPU但训练时内存暴涨崩溃问题求助
TensorFlow训练时GPU内存暴涨崩溃问题解决
问题现象
- 系统可正常识别NVIDIA GeForce GTX 960 4GB GPU,执行
tf.config.list_logical_devices('GPU')能输出正常设备信息 - 运行《AI and Machine Learning for Coders》示例代码时,调用
model.fit()方法卡在epoch 1/50,内存占用暴涨后程序崩溃 - 报错信息:
Process finished with exit code -1073740791 (0xC0000409),同时出现提示:I tensorflow/stream_executor/cuda/cuda_dnn.cc:384] Loaded cuDNN version 8401
环境配置
- tensorflow==2.9.0
- tensorflow-gpu==2.1.0
- CUDA v11.7
- CUDNN v8.4.1.50
测试代码
import tensorflow as tf data = tf.keras.datasets.fashion_mnist class myCallback(tf.keras.callbacks.Callback): def on_epoch_end(self, epoch, logs={}): if(logs.get('accuracy')>0.99): print("\nReached 99% accuracy so cancelling training!") self.model.stop_training = True callbacks = myCallback() (training_images, training_labels), (test_images, test_labels) = data.load_data() print(type(training_images)) training_images=training_images.reshape(60000, 28, 28, 1) training_images = training_images / 255.0 test_images = test_images.reshape(10000, 28, 28, 1) test_images = test_images / 255.0 model = tf.keras.models.Sequential([ tf.keras.layers.Conv2D(64, (3, 3), activation='relu', input_shape=(28, 28, 1)), tf.keras.layers.MaxPooling2D(2, 2), tf.keras.layers.Conv2D(64, (3, 3), activation='relu'), tf.keras.layers.MaxPooling2D(2,2), tf.keras.layers.Flatten(), tf.keras.layers.Dense(128, activation=tf.nn.relu), tf.keras.layers.Dense(10, activation=tf.nn.softmax)]) model.compile(optimizer='adam', loss='sparse_categorical_crossentropy', metrics=['accuracy']) model.summary() model.fit(training_images, training_labels, epochs=50, callbacks=[callbacks], verbose=1)
tf.config.list_logical_devices('GPU')输出
2022-09-10 23:30:04.052968: I tensorflow/core/platform/cpu_feature_guard.cc:193] This TensorFlow binary is optimized with oneAPI Deep Neural Network Library (oneDNN) to use the following CPU instructions in performance-critical operations: AVX AVX2 To enable them in other operations, rebuild TensorFlow with the appropriate compiler flags. 2022-09-10 23:30:04.496787: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1616] Created device /job:localhost/replica:0/task:0/device:GPU:0 with 2810 MB memory: -> device: 0, name: NVIDIA GeForce GTX 960, pci bus id: 0000:06:00.0, compute capability: 5.2 [LogicalDevice(name='/device:GPU:0', device_type='GPU')]
解决方案
1. 清理冲突的TensorFlow版本
同时安装tensorflow==2.9.0和tensorflow-gpu==2.1.0会引发版本冲突,这是核心问题之一:
- 卸载现有TensorFlow包:
pip uninstall tensorflow tensorflow-gpu -y - 重新安装单一版本:TensorFlow 2.4+之后,
tensorflow包已包含GPU支持,直接安装匹配版本:pip install tensorflow==2.9.0
2. 限制GPU内存按需分配
GTX 960仅4GB显存,默认TensorFlow会占用全部GPU内存,导致溢出。在导入tensorflow后立即添加以下代码:
gpus = tf.config.experimental.list_physical_devices('GPU') if gpus: try: for gpu in gpus: tf.config.experimental.set_memory_growth(gpu, True) logical_gpus = tf.config.experimental.list_logical_devices('GPU') print(len(gpus), "Physical GPUs,", len(logical_gpus), "Logical GPUs") except RuntimeError as e: print(e)
3. 调整训练批次大小
默认model.fit()的batch size为32,可适当减小以降低显存占用:
model.fit(training_images, training_labels, epochs=50, callbacks=[callbacks], verbose=1, batch_size=16)
4. 验证CUDA与TensorFlow兼容性
TensorFlow 2.9.0官方推荐搭配CUDA 11.2、cuDNN 8.1.0,当前CUDA 11.7虽可兼容,但如果问题仍存在,可降级CUDA到11.2版本,并确保cuDNN版本对应匹配。
内容的提问来源于stack exchange,提问作者9879ypxkj
相关产品推荐
相关产品推荐

