Ubuntu子系统中TensorFlow预测时Jupyter内核崩溃求助
WSL中TensorFlow运行GAN模型时Jupyter内核崩溃问题
环境信息
- 系统:Windows上的Linux子系统(WSL)
- TensorFlow版本:v2.12.0
- CUDA版本:11.3.58
- 已确认TensorFlow可连接GPU,但运行GAN模型执行图像预测步骤时Jupyter内核崩溃
终端错误日志
2023-03-29 10:40:36.194790: I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations. To enable the following instructions: AVX2 FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags. 2023-03-29 10:40:36.690448: W tensorflow/compiler/tf2tensorrt/utils/py_utils.cc:38] TF-TRT Warning: Could not find TensorRT 2023-03-29 10:40:37.609649: I tensorflow/compiler/xla/stream_executor/cuda/cuda_gpu_executor.cc:982] could not open file to read NUMA node: /sys/bus/pci/devices/0000:01:00.0/numa_node Your kernel may have been built without NUMA support. 2023-03-29 10:40:37.625212: I tensorflow/compiler/xla/stream_executor/cuda/cuda_gpu_executor.cc:982] could not open file to read NUMA node: /sys/bus/pci/devices/0000:01:00.0/numa_node Your kernel may have been built without NUMA support. 2023-03-29 10:40:37.625255: I tensorflow/compiler/xla/stream_executor/cuda/cuda_gpu_executor.cc:982] could not open file to read NUMA node: /sys/bus/pci/devices/0000:01:00.0/numa_node Your kernel may have been built without NUMA support. 2023-03-29 10:40:37.626900: I tensorflow/compiler/xla/stream_executor/cuda/cuda_gpu_executor.cc:982] could not open file to read NUMA node: /sys/bus/pci/devices/0000:01:00.0/numa_node Your kernel may have been built without NUMA support. 2023-03-29 10:40:37.626949: I tensorflow/compiler/xla/stream_executor/cuda/cuda_gpu_executor.cc:982] could not open file to read NUMA node: /sys/bus/pci/devices/0000:01:00.0/numa_node Your kernel may have been built without NUMA support. 2023-03-29 10:40:37.626979: I tensorflow/compiler/xla/stream_executor/cuda/cuda_gpu_executor.cc:982] could not open file to read NUMA node: /sys/bus/pci/devices/0000:01:00.0/numa_node Your kernel may have been built without NUMA support. 2023-03-29 10:40:38.273806: I tensorflow/compiler/xla/stream_executor/cuda/cuda_gpu_executor.cc:982] could not open file to read NUMA node: /sys/bus/pci/devices/0000:01:00.0/numa_node Your kernel may have been built without NUMA support. 2023-03-29 10:40:38.274180: I tensorflow/compiler/xla/stream_executor/cuda/cuda_gpu_executor.cc:982] could not open file to read NUMA node: /sys/bus/pci/devices/0000:01:00.0/numa_node Your kernel may have been built without NUMA support. 2023-03-29 10:40:38.274215: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1722] Could not identify NUMA node of platform GPU id 0, defaulting to 0. Your kernel may not have been built with NUMA support. 2023-03-29 10:40:38.274403: I tensorflow/compiler/xla/stream_executor/cuda/cuda_gpu_executor.cc:982] could not open file to read NUMA node: /sys/bus/pci/devices/0000:01:00.0/numa_node Your kernel may have been built without NUMA support. 2023-03-29 10:40:38.274534: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1635] Created device /job:localhost/replica:0/task:0/device:GPU:0 with 2042 MB memory: -> device: 0, name: Quadro T2000, pci bus id: 0000:01:00.0, compute capability: 7.5 2023-03-29 10:40:40.439728: I tensorflow/compiler/xla/stream_executor/cuda/cuda_dnn.cc:424] Loaded cuDNN version 8600 Could not load library libcudnn_cnn_train.so.8. Error: libcuda.so: cannot open shared object file: No such file or directory [I 10:40:43.009 NotebookApp] KernelRestarter: restarting kernel (1/5), keep random ports WARNING:root:kernel a03feebf-dd7a-41b4-9a9f-0333c160f338 restarted
已尝试操作
- 卸载并重新安装TensorFlow
- 所有组件通过Anaconda安装,怀疑路径配置存在问题
解决步骤
1. 修复libcuda.so缺失问题
WSL中TensorFlow找不到libcuda.so是常见问题,手动创建软链接即可:
打开WSL终端执行以下命令:
sudo ln -s /usr/lib/wsl/lib/libcuda.so.1 /usr/lib/x86_64-linux-gnu/libcuda.so sudo ln -s /usr/lib/wsl/lib/libcuda.so.1 /usr/lib/x86_64-linux-gnu/libcuda.so.1
该命令将WSL自带的GPU驱动库链接到系统默认搜索路径,让TensorFlow能找到它。
2. 验证Anaconda环境路径配置
- 激活你的Anaconda环境(替换为你自己的环境名,比如
tensorflow_env):
conda activate tensorflow_env
- 检查环境内的CUDA相关路径是否正确:
echo $LD_LIBRARY_PATH
正常输出应包含Anaconda环境的lib路径,比如~/anaconda3/envs/tensorflow_env/lib。如果没有,手动添加:
export LD_LIBRARY_PATH=$LD_LIBRARY_PATH:~/anaconda3/envs/tensorflow_env/lib
将这条命令写入~/.bashrc,确保每次打开终端自动生效:
echo 'export LD_LIBRARY_PATH=$LD_LIBRARY_PATH:~/anaconda3/envs/tensorflow_env/lib' >> ~/.bashrc source ~/.bashrc
3. 匹配TensorFlow与CUDA版本
TensorFlow 2.12.0官方推荐CUDA 11.8,当前使用的11.3版本不匹配可能导致崩溃:
- 卸载现有CUDA组件:
conda remove cuda cudnn
- 安装对应版本的CUDA和cudnn:
conda install cudatoolkit=11.8 cudnn=8.6.0 -c conda-forge pip install tensorflow==2.12.0
4. 限制GPU内存占用
GAN模型预测时可能因占用过多GPU内存导致崩溃,在代码开头添加以下内容限制内存按需分配:
import tensorflow as tf gpus = tf.config.list_physical_devices('GPU') if gpus: try: for gpu in gpus: tf.config.experimental.set_memory_growth(gpu, True) logical_gpus = tf.config.list_logical_devices('GPU') print(len(gpus), "Physical GPUs,", len(logical_gpus), "Logical GPUs") except RuntimeError as e: print(e)
内容的提问来源于stack exchange,提问作者ATunison
相关产品推荐
相关产品推荐

