Vertex AI Notebook中Cudnn版本不兼容致内核崩溃,如何升级?
解决Vertex AI中CuDNN与TensorFlow/CUDA不兼容的问题
方案1:使用Google官方预构建兼容镜像
Google提供了已配置好兼容版本的Vertex AI训练镜像,TensorFlow 2.10.0对应的GPU镜像预装了CUDA 11.2和CuDNN 8.1,完全匹配官方要求。
提交训练任务时指定该镜像即可,示例gcloud命令:
gcloud ai jobs submit training YOUR_JOB_NAME \ --region=YOUR_REGION \ --master-image-uri=gcr.io/deeplearning-platform-release/tf2-gpu.2-10 \ --scale-tier=BASIC_GPU \ --python-module=your.training.module \ --package-path=./your_package \ --project=YOUR_PROJECT_ID
方案2:自定义构建兼容镜像
若需基于现有镜像调整,可通过Dockerfile构建自定义镜像:
- 创建Dockerfile:
# 以TensorFlow 2.10 GPU官方镜像为基础 FROM tensorflow/tensorflow:2.10.0-gpu # 安装适配CUDA 11.2的CuDNN 8.1 RUN apt-get update && apt-get install -y --no-install-recommends \ libcudnn8=8.1.0.77-1+cuda11.2 \ libcudnn8-dev=8.1.0.77-1+cuda11.2 # 清理APT缓存以减小镜像体积 RUN apt-get clean && rm -rf /var/lib/apt/lists/*
- 构建并推送镜像至Google Container Registry:
docker build -t gcr.io/YOUR_PROJECT_ID/custom-tf210-cudnn81 . docker push gcr.io/YOUR_PROJECT_ID/custom-tf210-cudnn81
- 提交训练任务时替换镜像URI为上述自定义镜像地址即可。
验证版本兼容性
训练任务启动后,可在训练代码中加入以下片段确认版本:
import tensorflow as tf from tensorflow.python.platform import build_info as build print(f"tensorflow version: {tf.__version__}") print(f"Cuda Version: {build.build_info['cuda_version']}") print(f"Cudnn version: {build.build_info['cudnn_version']}")
确保输出中CuDNN版本显示为8.1。
内容的提问来源于stack exchange,提问作者santobedi
相关产品推荐
相关产品推荐

