基于nvidia/cuda镜像的自定义Docker容器无法使用GPU(已加--gpus all)
Docker容器GPU访问故障排查
问题现象
使用--gpus all启动容器后,nvidia-smi能正常识别GPU,但TensorFlow、PyTorch、ONNX Runtime均无法检测或使用GPU:
- 启动容器执行业务代码时,ONNX Runtime仅输出
CPUExecutionProvider - 直接执行
nvidia-smi则能正常显示GPU信息
测试命令
启动容器执行业务代码:
sudo docker run --gpus all mycontainer:latest
执行nvidia-smi验证GPU可见性:
sudo docker run --gpus all mycontainer:latest nvidia-smi
nvidia-smi输出
+-----------------------------------------------------------------------------+ | NVIDIA-SMI 495.29.05 Driver Version: 495.29.05 CUDA Version: 11.5 | |-------------------------------+----------------------+----------------------+ | GPU Name Persistence-M| Bus-Id Disp.A | Volatile Uncorr. ECC | | Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. | | | | MIG M. | |===============================+======================+======================| | 0 NVIDIA GeForce ... On | 00000000:01:00.0 Off | N/A | | N/A 44C P0 27W / N/A | 10MiB / 7982MiB | 0% Default | | | | N/A | +-------------------------------+----------------------+----------------------+ +-----------------------------------------------------------------------------+ | Processes: | | GPU GI CI PID Type Process name GPU Memory | | ID ID Usage | |=============================================================================| +-----------------------------------------------------------------------------+
容器Dockerfile
FROM nvidia/cuda:11.5.0-base-ubuntu20.04 WORKDIR /home COPY requirements.txt /home/requirements.txt # Add the deadsnakes PPA for Python 3.10 RUN apt-get update && \ apt-get install -y software-properties-common libgl1-mesa-glx cmake protobuf-compiler && \ add-apt-repository ppa:deadsnakes/ppa && \ apt-get update # Install Python 3.10 and dev packages RUN apt-get update && \ apt-get install -y python3.10 python3.10-dev python3-pip && \ rm -rf /var/lib/apt/lists/* # Install virtualenv RUN pip3 install virtualenv # Create a virtual environment with Python 3.10 RUN virtualenv -p python3.10 venv # Activate the virtual environment ENV PATH="/home/venv/bin:$PATH" # Install Python dependencies RUN pip3 install --upgrade pip \ && pip3 install --default-timeout=10000000 torch torchvision --extra-index-url https://download.pytorch.org/whl/cu116 \ && pip3 install --default-timeout=10000000 -r requirements.txt # Copy files COPY /src /home/src # Set the PYTHONPATH and LD_LIBRARY_PATH environment variable to include the CUDA libraries ENV PYTHONPATH=/usr/local/cuda-11.5/lib64 ENV LD_LIBRARY_PATH=/usr/local/cuda-11.5/lib64 # Set the CUDA_PATH and CUDA_HOME environment variable to point to the CUDA installation directory ENV CUDA_PATH=/usr/local/cuda-11.5 ENV CUDA_HOME=/usr/local/cuda-11.5 # Set the default command CMD ["sh", "-c", ". /home/venv/bin/activate && python main.py $@"]
问题原因及解决方法
1. PyTorch与CUDA版本不匹配
Docker基础镜像使用的是nvidia/cuda:11.5.0-base,但安装PyTorch时指定了cu116的索引源,导致PyTorch依赖的CUDA版本(11.6)与容器内CUDA runtime版本(11.5)冲突,无法加载GPU驱动。
解决:将PyTorch安装命令改为匹配CUDA 11.5的版本:
RUN pip3 install --upgrade pip \ && pip3 install --default-timeout=10000000 torch torchvision --extra-index-url https://download.pytorch.org/whl/cu115 \ && pip3 install --default-timeout=10000000 -r requirements.txt
2. PYTHONPATH设置错误
PYTHONPATH用于指定Python模块的查找路径,而当前设置成了CUDA的lib64目录,会干扰Python正常的模块搜索逻辑,导致GPU相关库无法被正确加载。
解决:删除错误的PYTHONPATH环境变量设置:
# 移除这一行 # ENV PYTHONPATH=/usr/local/cuda-11.5/lib64
3. ONNX Runtime默认安装CPU版本
默认pip install onnxruntime只会安装CPU版本,需要明确安装GPU版本才能支持CUDA。
解决:在requirements.txt中替换为GPU版本:
onnxruntime-gpu>=1.13.0 # 选择与CUDA 11.5兼容的版本
4. 验证修复效果
重新构建镜像后,进入容器执行以下命令验证:
- PyTorch GPU检测:
import torch print(torch.cuda.is_available()) # 应输出True print(torch.cuda.device_count()) # 应输出GPU数量 - ONNX Runtime GPU检测:
import onnxruntime as ort print(ort.get_available_providers()) # 应包含CUDAExecutionProvider
内容的提问来源于stack exchange,提问作者Moritz Müller
相关产品推荐
相关产品推荐

