能否通过CUDA Compat 12.1在CUDA驱动470上运行最新Triton Server?
问题
当前环境CUDA驱动版本为470.42.01,未进行更新,希望使用默认要求NVIDIA CUDA 12.1.0的Triton Inference Server 23.04。尝试通过以下Dockerfile构建镜像并运行:
FROM nvcr.io/nvidia/tritonserver:23.04-py3 COPY cuda-compat /cuda-compat RUN dpkg -i /cuda-compat/cuda-compat-12-1_530.30.02-1_amd64.deb RUN LD_LIBRARY_PATH="/usr/local/cuda-12.1/compat:${LD_LIBRARY_PATH}"
运行后出现错误,提示系统存在不支持的显示驱动/CUDA驱动组合(错误803),所有模型加载失败。不清楚是操作方法有误,还是该场景本身不支持?
错误信息如下:
ERROR: The NVIDIA Driver is present, but CUDA failed to initialize. GPU functionality will not be available. [[ System has unsupported display driver / cuda driver combination (error 803) ]] W0520 02:02:08.644624 1 pinned_memory_manager.cc:236] Unable to allocate pinned system memory, pinned memory pool will not be available: system has unsupported display driver / cuda driver combination E0520 02:02:08.644693 1 server.cc:230] Failed to initialize CUDA memory manager: unable to get number of CUDA devices: system has unsupported display driver / cuda driver combination W0520 02:02:08.644699 1 server.cc:237] failed to enable peer access for some device pairs I0520 02:02:08.650682 1 model_lifecycle.cc:459] loading: densenet_onnx:1 I0520 02:02:08.650712 1 model_lifecycle.cc:459] loading: inception_graphdef:1 I0520 02:02:08.650734 1 model_lifecycle.cc:459] loading: simple_int8:1 I0520 02:02:08.650753 1 model_lifecycle.cc:459] loading: simple_sequence:1 I0520 02:02:08.650771 1 model_lifecycle.cc:459] loading: simple:1 I0520 02:02:08.650791 1 model_lifecycle.cc:459] loading: simple_dyna_sequence:1 I0520 02:02:08.650812 1 model_lifecycle.cc:459] loading: simple_identity:1 I0520 02:02:08.650854 1 model_lifecycle.cc:459] loading: simple_string:1 I0520 02:02:08.651782 1 onnxruntime.cc:2504] TRITONBACKEND_Initialize: onnxruntime I0520 02:02:08.651801 1 onnxruntime.cc:2514] Triton TRITONBACKEND API version: 1.12 I0520 02:02:08.651806 1 onnxruntime.cc:2520] 'onnxruntime' TRITONBACKEND API version: 1.12 I0520 02:02:08.651811 1 onnxruntime.cc:2550] backend configuration: {\"cmdline\":{\"auto-complete-config\":\"true\",\"backend-directory\":\"/opt/tritonserver/backends\",\"min-compute-capability\":\"6.000000\",\"default-max-batch-size\":\"4\"}} E0520 02:02:08.663226 1 model_lifecycle.cc:597] failed to load 'densenet_onnx' version 1: Internal: unable to get number of CUDA devices: system has unsupported display driver / cuda driver combination I0520 02:02:08.883231 1 tensorflow.cc:2565] TRITONBACKEND_Initialize: tensorflow I0520 02:02:08.883260 1 tensorflow.cc:2575] Triton TRITONBACKEND API version: 1.12 I0520 02:02:08.883265 1 tensorflow.cc:2581] 'tensorflow' TRITONBACKEND API version: 1.12 I0520 02:02:08.883269 1 tensorflow.cc:2605] backend configuration: {\"cmdline\":{\"auto-complete-config\":\"true\",\"backend-directory\":\"/opt/tritonserver/backends\",\"min-compute-capability\":\"6.000000\",\"default-max-batch-size\":\"4\"}} E0520 02:02:08.883301 1 model_lifecycle.cc:597] failed to load 'simple_int8' version 1: Internal: unable to get number of CUDA devices: system has unsupported display driver / cuda driver combination E0520 02:02:08.883308 1 model_lifecycle.cc:597] failed to load 'inception_graphdef' version 1: Internal: unable to get number of CUDA devices: system has unsupported display driver / cuda driver combination E0520 02:02:08.883314 1 model_lifecycle.cc:597] failed to load 'simple' version 1: Internal: unable to get number of CUDA devices: system has unsupported display driver / cuda driver combination E0520 02:02:08.883334 1 model_lifecycle.cc:597] failed to load 'simple_identity' version 1: Internal: unable to get number of CUDA devices: system has unsupported display driver / cuda driver combination E0520 02:02:08.883333 1 model_lifecycle.cc:597] failed to load 'simple_sequence' version 1: Internal: unable to get number of CUDA devices: system has unsupported display driver / cuda driver combination E0520 02:02:08.883347 1 model_lifecycle.cc:597] failed to load 'simple_dyna_sequence' version 1: Internal: unable to get number of CUDA devices: system has unsupported display driver / cuda driver combination E0520 02:02:08.883365 1 model_lifecycle.cc:597] failed to load 'simple_string' version 1: Internal: unable to get number of CUDA devices: system has unsupported display driver / cuda driver combination I0520 02:02:08.883525 1 server.cc:610] +-------------+-----------------------------------------------------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Backend | Path | Config | +-------------+-----------------------------------------------------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------+ | onnxruntime | /opt/tritonserver/backends/onnxruntime/libtriton_onnxruntime.so | {\"cmdline\":{\"auto-complete-config\":\"true\",\"backend-directory\":\"/opt/tritonserver/backends\",\"min-compute-capability\":\"6.000000\",\"default-max-batch-size\":\"4\"}} | | tensorflow | /opt/tritonserver/backends/tensorflow/libtriton_tensorflow.so | {\"cmdline\":{\"auto-complete-config\":\"true\",\"backend-directory\":\"/opt/tritonserver/backends\",\"min-compute-capability\":\"6.000000\",\"default-max-batch-size\":\"4\"}} | +-------------+-----------------------------------------------------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------+ I0520 02:02:08.883603 1 server.cc:653] +----------------------+---------+------------------------------------------------------------------------------------------------------------------------------+ | Model | Version | Status | +----------------------+---------+------------------------------------------------------------------------------------------------------------------------------+ | densenet_onnx | 1 | UNAVAILABLE: Internal: unable to get number of CUDA devices: system has unsupported display driver / cuda driver combination | | inception_graphdef | 1 | UNAVAILABLE: Internal: unable to get number of CUDA devices: system has unsupported display driver / cuda driver combination | +----------------------+---------+------------------------------------------------------------------------------------------------------------------------------+ W0520 02:02:08.919379 1 metrics.cc:792] Cannot get CUDA device count, GPU metrics will not be available I0520 02:02:08.919635 1 metrics.cc:701] Collecting CPU metrics I0520 02:02:08.919802 1 server.cc:284] Waiting for in-flight requests to complete. I0520 02:02:08.919810 1 server.cc:300] Timeout 30: Found 0 model versions that have in-flight inferences I0520 02:02:08.919817 1 server.cc:315] All models are stopped, unloading models I0520 02:02:08.919823 1 server.cc:322] Timeout 30: Found 0 live models and 0 in-flight non-inference requests error: creating server: Internal - failed to load all models
解决方案
核心原因
CUDA Compat包的作用是让旧版本主机驱动兼容容器内的高版本CUDA Runtime,但存在严格的版本匹配限制:
- 你的主机驱动是470.42.01,属于CUDA 11.4系列,支持的最高CUDA Runtime版本为11.7
cuda-compat-12-1包要求的最低主机驱动版本是525.60.13(对应CUDA 12.0及以上),470.x驱动版本远低于这个要求,因此无法兼容,导致错误803。
可行解决方法
方法1:降级Triton版本(无需升级驱动)
选择与470.x驱动兼容的Triton镜像,这类镜像基于CUDA 11.7及以下版本。例如:
nvcr.io/nvidia/tritonserver:22.11-py3
该版本Triton对应CUDA 11.7,完全适配470.x驱动,直接使用官方镜像即可正常运行,无需额外修改。
方法2:升级主机NVIDIA驱动(保留Triton 23.04)
若必须使用Triton 23.04,需将主机的NVIDIA驱动升级到525.60.13或更高版本(CUDA 12.1的最低驱动要求)。升级完成后,直接使用官方Triton 23.04镜像即可,无需安装CUDA Compat包。
你的Dockerfile存在的小问题
你在构建镜像时用RUN指令设置LD_LIBRARY_PATH,这个环境变量仅在当前构建步骤生效,容器运行时不会保留。正确的方式应该使用ENV指令:
ENV LD_LIBRARY_PATH="/usr/local/cuda-12.1/compat:${LD_LIBRARY_PATH}"
但即使修正这一点,也解决不了根本的版本兼容问题,因为驱动版本不符合要求。
内容的提问来源于stack exchange,提问作者聂小涛
相关产品推荐
相关产品推荐

