TensorFlow训练过程中崩溃求助(WSL2 Ubuntu 22.04环境)
TensorFlow训练崩溃问题解决方案
问题现象
执行以下训练代码时,运行至Epoch 1/25阶段直接崩溃:
model.fit(X, y,batch_size=5,epochs=25, validation_split=0.3,callbacks=[tensorboard])
崩溃前的日志输出:
2023-06-14 12:53:37.515906: I tensorflow/compiler/xla/stream_executor/cuda/cuda_gpu_executor.cc:982] could not open file to read NUMA node: /sys/bus/pci/devices/0000:06:00.0/numa_node Your kernel may have been built without NUMA support. 2023-06-14 12:53:37.516050: I tensorflow/compiler/xla/stream_executor/cuda/cuda_gpu_executor.cc:982] could not open file to read NUMA node: /sys/bus/pci/devices/0000:06:00.0/numa_node Your kernel may have been built without NUMA support. 2023-06-14 12:53:37.516128: I tensorflow/compiler/xla/stream_executor/cuda/cuda_gpu_executor.cc:982] could not open file to read NUMA node: /sys/bus/pci/devices/0000:06:00.0/numa_node Your kernel may have been built without NUMA support. 2023-06-14 12:53:38.588014: I tensorflow/compiler/xla/stream_executor/cuda/cuda_gpu_executor.cc:982] could not open file to read NUMA node: /sys/bus/pci/devices/0000:06:00.0/numa_node Your kernel may have been built without NUMA support. 2023-06-14 12:53:38.588102: I tensorflow/compiler/xla/stream_executor/cuda/cuda_gpu_executor.cc:982] could not open file to read NUMA node: /sys/bus/pci/devices/0000:06:00.0/numa_node Your kernel may have been built without NUMA support. 2023-06-14 12:53:38.588112: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1722] Could not identify NUMA node of platform GPU id 0, defaulting to 0. Your kernel may not have been built with NUMA support. 2023-06-14 12:53:38.588161: I tensorflow/compiler/xla/stream_executor/cuda/cuda_gpu_executor.cc:982] could not open file to read NUMA node: /sys/bus/pci/devices/0000:06:00.0/numa_node Your kernel may have been built without NUMA support. 2023-06-14 12:53:38.588209: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1635] Created device /job:localhost/replica:0/task:0/device:GPU:0 with 5463 MB memory: -> device: 0, name: NVIDIA GeForce RTX 2060 SUPER, pci bus id: 0000:06:00.0, compute capability: 7.5 Epoch 1/25
已确认的环境信息
- 系统:WSL 2 + Ubuntu 22.04
- TensorFlow版本:2.12.0
- 设备检测代码及输出:
import tensorflow as tf from tensorflow.python.client import device_lib print("devices: ", [d.name for d in device_lib.list_local_devices()]) print("GPUs: ", tf.config.list_physical_devices('GPU')) print("TF v.: ", tf.__version__)
devices: ['/device:CPU:0', '/device:GPU:0'] GPUs: [PhysicalDevice(name='/physical_device:GPU:0', device_type='GPU')] TF v.: 2.12.0
nvidia-smi输出:
+---------------------------------------------------------------------------------------+ | NVIDIA-SMI 535.43.02 Driver Version: 535.98 CUDA Version: 12.2 | |-----------------------------------------+----------------------+----------------------+| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC | | Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. | | | | MIG M. | |=========================================+======================+======================| | 0 NVIDIA GeForce RTX 2060 ... On | 00000000:06:00.0 On | N/A | | 55% 45C P8 15W / 184W | 7645MiB / 8192MiB | 6% Default | | | | N/A | +-----------------------------------------+----------------------+----------------------+ +---------------------------------------------------------------------------------------+ | Processes: | | GPU GI CI PID Type Process name GPU Memory | | ID ID Usage | |=======================================================================================| | 0 N/A N/A 23 G /Xwayland N/A | | 0 N/A N/A 894 C /python3.9 N/A | +---------------------------------------------------------------------------------------+
已尝试重装TensorFlow和CUDA,问题未解决。
解决步骤
1. 限制GPU显存使用
从nvidia-smi输出可见,GPU已占用7645MiB,剩余显存不足是崩溃核心原因。在训练代码开头添加以下代码调整显存分配策略:
gpus = tf.config.list_physical_devices('GPU') if gpus: try: # 按需分配显存,避免一次性占满 for gpu in gpus: tf.config.experimental.set_memory_growth(gpu, True) # 也可固定分配显存,比如限制为4GB # tf.config.set_logical_device_configuration( # gpus[0], # [tf.config.LogicalDeviceConfiguration(memory_limit=4096)] # ) except RuntimeError as e: print(e)
2. 调小batch_size
当前batch_size=5仍可能导致显存溢出,尝试将batch_size改为2或3,减少单次迭代的显存占用。
3. 消除NUMA警告(非崩溃原因,仅优化日志)
NUMA警告是WSL2内核特性问题,不影响功能。若要屏蔽该类INFO日志,启动脚本前设置环境变量:
export TF_CPP_MIN_LOG_LEVEL=2
4. 匹配TensorFlow与CUDA版本
TensorFlow 2.12.0官方推荐CUDA 11.8,当前系统安装的CUDA 12.2版本不兼容可能引发隐性崩溃,操作步骤:
- 卸载现有CUDA:
sudo apt-get --purge remove "*cuda*" "*cudnn*" sudo rm -rf /usr/local/cuda*
- 安装CUDA 11.8:
wget https://developer.download.nvidia.com/compute/cuda/11.8.0/local_installers/cuda_11.8.0_520.61.05_linux.run sudo sh cuda_11.8.0_520.61.05_linux.run --override
- 配置环境变量(添加至
~/.bashrc):
export PATH=/usr/local/cuda-11.8/bin${PATH:+:${PATH}} export LD_LIBRARY_PATH=/usr/local/cuda-11.8/lib64${LD_LIBRARY_PATH:+:${LD_LIBRARY_PATH}}
- 安装对应cuDNN 8.6.0:下载cuDNN 8.6.0 for CUDA 11.x的tar包,解压后复制文件:
tar -xzvf cudnn-linux-x86_64-8.6.0.163_cuda11-archive.tar.xz sudo cp cudnn-linux-x86_64-8.6.0.163_cuda11-archive/include/cudnn*.h /usr/local/cuda-11.8/include sudo cp cudnn-linux-x86_64-8.6.0.163_cuda11-archive/lib/libcudnn* /usr/local/cuda-11.8/lib64 sudo chmod a+r /usr/local/cuda-11.8/include/cudnn*.h /usr/local/cuda-11.8/lib64/libcudnn*
5. 释放图形界面占用的显存
nvidia-smi显示/Xwayland占用GPU资源,若无需WSL2图形界面,可关闭该进程释放显存:
pkill Xwayland
内容的提问来源于stack exchange,提问作者proto
相关产品推荐
相关产品推荐

