运行TensorFlow GPU检测代码时遇cuInit调用失败错误求助
CUDA_ERROR_UNKNOWN问题 问题重现
运行以下GPU检测代码:
import tensorflow as tf if tf.config.list_physical_devices('GPU'): print('GPU is available and configured for TensorFlow') else: print('GPU is not available or not properly configured for TensorFlow')
出现错误日志:
2023-04-17 11:43:09.186983: E tensorflow/stream_executor/cuda/cuda_driver.cc:265] failed call to cuInit: CUDA_ERROR_UNKNOWN: unknown error
控制台输出:
GPU is not available or not properly configured for TensorFlow
解决方案
验证NVIDIA驱动与CUDA版本兼容性
确认当前NVIDIA驱动版本和所用TensorFlow要求的CUDA版本匹配。通过nvidia-smi命令查看驱动版本,对照TensorFlow版本对应的依赖要求调整,比如TensorFlow 2.15对应CUDA 12.2、cuDNN 8.9。重启相关进程与服务
关闭所有TensorFlow相关进程,Windows用户可在服务列表中重启NVIDIA Display Container等相关服务;Linux用户执行sudo systemctl restart nvidia-persistenced,之后重新运行检测代码。检查CUDA工具包安装与环境变量
运行nvcc --version验证CUDA编译器是否正常工作。若命令不存在,需将CUDA的bin和libnvvp路径加入系统PATH;Linux系统还要确保LD_LIBRARY_PATH包含CUDA的lib64目录。排查GPU硬件状态
用nvidia-smi查看GPU是否被其他进程占用,或存在硬件异常。若GPU使用率100%,先终止占用进程,或重启机器后再尝试。重新安装匹配的依赖包
卸载当前TensorFlow(pip uninstall tensorflow),按官方要求重新安装对应版本的CUDA、cuDNN,再安装匹配的TensorFlow GPU版本(如pip install tensorflow==2.15.0)。
内容的提问来源于stack exchange,提问作者Ghazaleh Alizadegan

