Docker GPU镜像报错:内核版本535.129.3与DSO版本545.23.6不匹配
我原本以为带GPU支持的Docker镜像能解决CUDA安装问题,执行了以下命令:
docker pull tensorflow/tensorflow:latest-gpu-jupyter docker run --gpus all -it --rm -p 8889:8888 tensorflow/tensorflow:latest-gpu-jupyter
但在检查Jupyter服务器的GPU支持情况时,出现了CUDA版本不匹配的错误,核心问题是主机系统的CUDA内核版本(535.129.3)与镜像内的CUDA DSO版本(545.23.6)不一致,完整报错信息如下:
I tensorflow/core/util/port.cc:113] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable
TF_ENABLE_ONEDNN_OPTS=0.
E external/local_xla/xla/stream_executor/cuda/cuda_dnn.cc:9261] Unable to register cuDNN factory: Attempting to register factory for plugin cuDNN when one has already been registered
E external/local_xla/xla/stream_executor/cuda/cuda_fft.cc:607] Unable to register cuFFT factory: Attempting to register factory for plugin cuFFT when one has already been registered
E external/local_xla/xla/stream_executor/cuda/cuda_blas.cc:1515] Unable to register cuBLAS factory: Attempting to register factory for plugin cuBLAS when one has already been registered
I tensorflow/core/platform/cpu_feature_guard.cc:182] This TensorFlow binary is optimized to use available CPU instructions in performance-critical operations.
To enable the following instructions: AVX2 AVX_VNNI FMA, in other operations, rebuild TensorFlow with the appropriate compiler flags.
E external/local_xla/xla/stream_executor/cuda/cuda_driver.cc:274] failed call to cuInit: CUDA_ERROR_COMPAT_NOT_SUPPORTED_ON_DEVICE: forward compatibility was attempted on non supported HW
I external/local_xla/xla/stream_executor/cuda/cuda_diagnostics.cc:129] retrieving CUDA diagnostic information for host: 0e940b862ceb
I external/local_xla/xla/stream_executor/cuda/cuda_diagnostics.cc:136] hostname: 0e940b862ceb
I external/local_xla/xla/stream_executor/cuda/cuda_diagnostics.cc:159] libcuda reported version is: 545.23.6
I external/local_xla/xla/stream_executor/cuda/cuda_diagnostics.cc:163] kernel reported version is: 535.129.3
E external/local_xla/xla/stream_executor/cuda/cuda_diagnostics.cc:244] kernel version 535.129.3 does not match DSO version 545.23.6 -- cannot find working devices in this configuration
解决方法
1. 升级主机CUDA驱动版本
把主机上的CUDA驱动升级到和镜像内DSO版本(545.23.6)一致或更高的兼容版本,确保内核驱动与用户空间库版本匹配。
2. 使用和主机驱动兼容的TensorFlow镜像
放弃使用latest-gpu-jupyter通用标签,选择与主机CUDA驱动版本匹配的特定镜像:
- 先在主机上运行
nvidia-smi,查看输出里的CUDA Version字段,确认驱动支持的最高CUDA toolkit版本 - 选择对应版本的TensorFlow GPU镜像,比如如果主机支持CUDA 12.2,可执行以下命令:
docker pull tensorflow/tensorflow:2.15.0-gpu-jupyter-cuda12.2 docker run --gpus all -it --rm -p 8889:8888 tensorflow/tensorflow:2.15.0-gpu-jupyter-cuda12.2
3. 强制启用CUDA向前兼容(仅部分硬件支持)
如果你的GPU硬件支持向前兼容,可尝试添加环境变量绕过版本检查,但这种方法可能存在稳定性风险:
docker run --gpus all -it --rm -p 8889:8888 -e NVIDIA_DISABLE_REQUIRE=1 tensorflow/tensorflow:latest-gpu-jupyter
内容的提问来源于stack exchange,提问作者Endre Moen

