You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow训练过程中崩溃求助(WSL2 Ubuntu 22.04环境)

TensorFlow训练崩溃问题解决方案

问题现象

执行以下训练代码时,运行至Epoch 1/25阶段直接崩溃:

model.fit(X, y,batch_size=5,epochs=25, validation_split=0.3,callbacks=[tensorboard])

崩溃前的日志输出:

2023-06-14 12:53:37.515906: I tensorflow/compiler/xla/stream_executor/cuda/cuda_gpu_executor.cc:982] could not open file to read NUMA node: /sys/bus/pci/devices/0000:06:00.0/numa_node
Your kernel may have been built without NUMA support.
2023-06-14 12:53:37.516050: I tensorflow/compiler/xla/stream_executor/cuda/cuda_gpu_executor.cc:982] could not open file to read NUMA node: /sys/bus/pci/devices/0000:06:00.0/numa_node
Your kernel may have been built without NUMA support.
2023-06-14 12:53:37.516128: I tensorflow/compiler/xla/stream_executor/cuda/cuda_gpu_executor.cc:982] could not open file to read NUMA node: /sys/bus/pci/devices/0000:06:00.0/numa_node
Your kernel may have been built without NUMA support.
2023-06-14 12:53:38.588014: I tensorflow/compiler/xla/stream_executor/cuda/cuda_gpu_executor.cc:982] could not open file to read NUMA node: /sys/bus/pci/devices/0000:06:00.0/numa_node
Your kernel may have been built without NUMA support.
2023-06-14 12:53:38.588102: I tensorflow/compiler/xla/stream_executor/cuda/cuda_gpu_executor.cc:982] could not open file to read NUMA node: /sys/bus/pci/devices/0000:06:00.0/numa_node
Your kernel may have been built without NUMA support.
2023-06-14 12:53:38.588112: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1722] Could not identify NUMA node of platform GPU id 0, defaulting to 0.  Your kernel may not have been built with NUMA support.
2023-06-14 12:53:38.588161: I tensorflow/compiler/xla/stream_executor/cuda/cuda_gpu_executor.cc:982] could not open file to read NUMA node: /sys/bus/pci/devices/0000:06:00.0/numa_node
Your kernel may have been built without NUMA support.
2023-06-14 12:53:38.588209: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1635] Created device /job:localhost/replica:0/task:0/device:GPU:0 with 5463 MB memory:  -> device: 0, name: NVIDIA GeForce RTX 2060 SUPER, pci bus id: 0000:06:00.0, compute capability: 7.5
Epoch 1/25

已确认的环境信息

  • 系统:WSL 2 + Ubuntu 22.04
  • TensorFlow版本:2.12.0
  • 设备检测代码及输出:
import tensorflow as tf
from tensorflow.python.client import device_lib

print("devices: ", [d.name for d in device_lib.list_local_devices()])
print("GPUs:    ", tf.config.list_physical_devices('GPU'))
print("TF v.:   ", tf.__version__)
devices:  ['/device:CPU:0', '/device:GPU:0']
GPUs:     [PhysicalDevice(name='/physical_device:GPU:0', device_type='GPU')]
TF v.:    2.12.0
  • nvidia-smi输出:
+---------------------------------------------------------------------------------------+
| NVIDIA-SMI 535.43.02              Driver Version: 535.98       CUDA Version: 12.2     |
|-----------------------------------------+----------------------+----------------------+| GPU  Name                 Persistence-M | Bus-Id        Disp.A | Volatile Uncorr. ECC |
| Fan  Temp   Perf          Pwr:Usage/Cap |         Memory-Usage | GPU-Util  Compute M. |
|                                         |                      |               MIG M. |
|=========================================+======================+======================|
|   0  NVIDIA GeForce RTX 2060 ...    On  | 00000000:06:00.0  On |                  N/A |
| 55%   45C    P8              15W / 184W |   7645MiB /  8192MiB |      6%      Default |
|                                         |                      |                  N/A |
+-----------------------------------------+----------------------+----------------------+

+---------------------------------------------------------------------------------------+
| Processes:                                                                            |
|  GPU   GI   CI        PID   Type   Process name                            GPU Memory |
|        ID   ID                                                             Usage      |
|=======================================================================================|
|    0   N/A  N/A        23      G   /Xwayland                                 N/A      |
|    0   N/A  N/A       894      C   /python3.9                                N/A      |
+---------------------------------------------------------------------------------------+

已尝试重装TensorFlow和CUDA,问题未解决。

解决步骤

1. 限制GPU显存使用

从nvidia-smi输出可见,GPU已占用7645MiB,剩余显存不足是崩溃核心原因。在训练代码开头添加以下代码调整显存分配策略:

gpus = tf.config.list_physical_devices('GPU')
if gpus:
    try:
        # 按需分配显存,避免一次性占满
        for gpu in gpus:
            tf.config.experimental.set_memory_growth(gpu, True)
        # 也可固定分配显存,比如限制为4GB
        # tf.config.set_logical_device_configuration(
        #     gpus[0],
        #     [tf.config.LogicalDeviceConfiguration(memory_limit=4096)]
        # )
    except RuntimeError as e:
        print(e)

2. 调小batch_size

当前batch_size=5仍可能导致显存溢出,尝试将batch_size改为2或3,减少单次迭代的显存占用。

3. 消除NUMA警告(非崩溃原因,仅优化日志)

NUMA警告是WSL2内核特性问题,不影响功能。若要屏蔽该类INFO日志,启动脚本前设置环境变量:

export TF_CPP_MIN_LOG_LEVEL=2

4. 匹配TensorFlow与CUDA版本

TensorFlow 2.12.0官方推荐CUDA 11.8,当前系统安装的CUDA 12.2版本不兼容可能引发隐性崩溃,操作步骤:

  • 卸载现有CUDA:
sudo apt-get --purge remove "*cuda*" "*cudnn*"
sudo rm -rf /usr/local/cuda*
  • 安装CUDA 11.8:
wget https://developer.download.nvidia.com/compute/cuda/11.8.0/local_installers/cuda_11.8.0_520.61.05_linux.run
sudo sh cuda_11.8.0_520.61.05_linux.run --override
  • 配置环境变量(添加至~/.bashrc):
export PATH=/usr/local/cuda-11.8/bin${PATH:+:${PATH}}
export LD_LIBRARY_PATH=/usr/local/cuda-11.8/lib64${LD_LIBRARY_PATH:+:${LD_LIBRARY_PATH}}
  • 安装对应cuDNN 8.6.0:下载cuDNN 8.6.0 for CUDA 11.x的tar包,解压后复制文件:
tar -xzvf cudnn-linux-x86_64-8.6.0.163_cuda11-archive.tar.xz
sudo cp cudnn-linux-x86_64-8.6.0.163_cuda11-archive/include/cudnn*.h /usr/local/cuda-11.8/include
sudo cp cudnn-linux-x86_64-8.6.0.163_cuda11-archive/lib/libcudnn* /usr/local/cuda-11.8/lib64
sudo chmod a+r /usr/local/cuda-11.8/include/cudnn*.h /usr/local/cuda-11.8/lib64/libcudnn*

5. 释放图形界面占用的显存

nvidia-smi显示/Xwayland占用GPU资源,若无需WSL2图形界面,可关闭该进程释放显存:

pkill Xwayland

内容的提问来源于stack exchange,提问作者proto

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 22:39:59