Azure机器学习GPU实例CuDNN版本不兼容问题的解决方案咨询
Azure ML Studio GPU虚拟机Keras训练CuDNN版本冲突解决
问题场景
在Microsoft Azure Machine Learning Studio的GPU虚拟机笔记本中训练Keras模型时,触发CuDNN版本不兼容错误,错误日志如下:
2023-04-27 09:56:21.098249: E tensorflow/compiler/xla/stream_executor/cuda/cuda_dnn.cc:417] Loaded runtime CuDNN library: 8.2.4 but source was compiled with: 8.6.0. CuDNN library needs to have matching major version and equal or higher minor version. If using a binary install, upgrade your CuDNN library. If building from sources, make sure the library loaded at runtime is compatible with the version specified during compile configuration. 2023-04-27 09:56:21.099011: W tensorflow/core/framework/op_kernel.cc:1830] OP_REQUIRES failed at pooling_ops_common.cc:412 : UNIMPLEMENTED: DNN library is not found. 2023-04-27 09:56:21.099050: I tensorflow/core/common_runtime/executor.cc:1197] [/job:localhost/replica:0/task:0/device:GPU:0] (DEBUG INFO) Executor start aborting (this does not indicate an error and you can ignore this message): UNIMPLEMENTED: DNN library is not found. [[{{node model_2/max_pooling1d_6/MaxPool}}]] 2023-04-27 09:56:21.100704: E tensorflow/compiler/xla/stream_executor/cuda/cuda_dnn.cc:417] Loaded runtime CuDNN library: 8.2.4 but source was compiled with: 8.6.0. CuDNN library needs to have matching major version and equal or higher minor version. If using a binary install, upgrade your CuDNN library. If building from sources, make sure the library loaded at runtime is compatible with the version specified during compile configuration. 2023-04-27 09:56:21.101366: W tensorflow/core/framework/op_kernel.cc:1830] OP_REQUIRES failed at pooling_ops_common.cc:412 : UNIMPLEMENTED: DNN library is not found.
解决方案
1. 使用Azure ML预配置GPU环境
直接选用Azure ML提供的TensorFlow官方预配置镜像,这类镜像已预先适配好兼容的CUDA、CuDNN和TensorFlow版本,彻底规避手动配置的版本冲突:
- 创建新笔记本实例时,在环境选择步骤中挑选对应版本的TensorFlow GPU镜像(如
TensorFlow 2.12 GPU); - 现有笔记本可通过顶部环境切换菜单,切换到预配置的TensorFlow GPU环境。
2. 手动升级CuDNN(自定义环境适用)
针对已自定义的环境,按以下步骤升级CuDNN至匹配版本(需与TensorFlow编译版本一致,此处为8.6.0+):
- 卸载旧版本CuDNN:
sudo apt-get remove --purge libcudnn8 - 添加NVIDIA官方软件源(若未配置):
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2004/x86_64/cuda-keyring_1.0-1_all.deb sudo dpkg -i cuda-keyring_1.0-1_all.deb sudo apt-get update - 安装指定版本CuDNN(需匹配当前CUDA版本,示例为CUDA 11.8):
sudo apt-get install libcudnn8=8.6.0.163-1+cuda11.8 - 验证安装结果:
确认输出版本为8.6.0及以上。dpkg -l | grep libcudnn8
3. 通过Conda管理兼容环境
利用Conda创建独立环境,强制指定兼容的版本组合:
- 创建并激活新环境:
conda create -n tf-compat-env tensorflow-gpu==2.12.0 cudnn==8.6.0 cudatoolkit==11.8 -y conda activate tf-compat-env - 在该环境中运行Keras训练代码,版本冲突问题即可解决。
内容的提问来源于stack exchange,提问作者Gideon Kogan
相关产品推荐
相关产品推荐

