Weaviate ResNet50容器NVIDIA驱动及CUDA不可用问题求助
解决Weaviate ResNet50 PyTorch容器CUDA不可用问题
问题描述
部署Weaviate ResNet50的PyTorch版本容器时,出现NVIDIA驱动未找到的错误:
weaviate-i2v-neural-1 | INFO: Started server process [7] weaviate-i2v-neural-1 | INFO: Waiting for application startup. weaviate-i2v-neural-1 | INFO: CUDA_CORE set to cuda:0 weaviate-i2v-neural-1 | /usr/local/lib/python3.11/site-packages/torchvision/models/_utils.py:208: UserWarning: The parameter 'pretrained' is deprecated since 0.13 and may be removed in the future, please use 'weights' instead. weaviate-i2v-neural-1 | warnings.warn( weaviate-i2v-neural-1 | /usr/local/lib/python3.11/site-packages/torchvision/models/_utils.py:223: UserWarning: Arguments other than a weight enum or `None` for 'weights' are deprecated since 0.13 and may be removed in the future. The current behavior is equivalent to passing `weights=ResNet50_Weights.IMAGENET1K_V1`. You can also use `weights=ResNet50_Weights.DEFAULT` to get the most up-to-date weights. weaviate-i2v-neural-1 | warnings.warn(msg) weaviate-i2v-neural-1 | ERROR: Traceback (most recent call last): weaviate-i2v-neural-1 | File "/usr/local/lib/python3.11/site-packages/starlette/routing.py", line 677, in lifespan weaviate-i2v-neural-1 | async with self.lifespan_context(app) as maybe_state: weaviate-i2v-neural-1 | File "/usr/local/lib/python3.11/site-packages/starlette/routing.py", line 566, in __aenter__ weaviate-i2v-neural-1 | await self._router.startup() weaviate-i2v-neural-1 | File "/usr/local/lib/python3.11/site-packages/starlette/routing.py", line 656, in startup weaviate-i2v-neural-1 | handler() weaviate-i2v-neural-1 | File "/app/app.py", line 29, in startup_event weaviate-i2v-neural-1 | imgVec = ImageVectorizer(cuda_support, cuda_core) weaviate-i2v-neural-1 | ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ weaviate-i2v-neural-1 | File "/app/vectorizer.py", line 13, in __init__ weaviate-i2v-neural-1 | self.img2vec = Img2VecPytorch(cuda_support, cuda_core) weaviate-i2v-neural-1 | ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ weaviate-i2v-neural-1 | File "/app/image2vec.py", line 16, in __init__ weaviate-i2v-neural-1 | self.model = self.model.to(self.device) weaviate-i2v-neural-1 | ^^^^^^^^^^^^^^^^^^^^^^^^^^ weaviate-i2v-neural-1 | File "/usr/local/lib/python3.11/site-packages/torch/nn/modules/module.py", line 1145, in to weaviate-i2v-neural-1 | return self._apply(convert) weaviate-i2v-neural-1 | ^^^^^^^^^^^^^^^^^^^^ weaviate-i2v-neural-1 | File "/usr/local/lib/python3.11/site-packages/torch/nn/modules/module.py", line 797, in _apply weaviate-i2v-neural-1 | module._apply(fn) weaviate-i2v-neural-1 | File "/usr/local/lib/python3.11/site-packages/torch/nn/modules/module.py", line 820, in _apply weaviate-i2v-neural-1 | param_applied = fn(param) weaviate-i2v-neural-1 | ^^^^^^^^^ weaviate-i2v-neural-1 | File "/usr/local/lib/python3.11/site-packages/torch/nn/modules/module.py", line 1143, in convert weaviate-i2v-neural-1 | return t.to(device, dtype if t.is_floating_point() or t.is_complex() else None, non_blocking) weaviate-i2v-neural-1 | ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ weaviate-i2v-neural-1 | File "/usr/local/lib/python3.11/site-packages/torch/cuda/__init__.py", line 247, in _lazy_init weaviate-i2v-neural-1 | torch._C._cuda_init() weaviate-i2v-neural-1 | RuntimeError: Found no NVIDIA driver on your system. Please check that you have an NVIDIA GPU and installed a driver from http://www.nvidia.com/Download/index.aspx weaviate-i2v-neural-1 | weaviate-i2v-neural-1 | ERROR: Application startup failed. Exiting. weaviate-i2v-neural-1 exited with code 3
设备为RTX3060,其他机器学习项目可正常运行。添加以下Docker Compose配置后,主机nvidia-smi可识别GPU,但容器仍报错RuntimeError: No CUDA GPUs are available:
i2v-neural: image: semitechnologies/img2vec-keras:resnet50 environment: ENABLE_CUDA: '1' CUDA_VISIBLE_DEVICES: 'all' deploy: resources: reservations: devices: - driver: nvidia count: 1 capabilities: [gpu]
解决方案
修正容器镜像版本
你要部署的是PyTorch版本的ResNet50向量器,但配置中使用了Keras版本的镜像semitechnologies/img2vec-keras:resnet50,这会导致环境不匹配。替换为PyTorch版本镜像:image: semitechnologies/img2vec-pytorch:resnet50验证Docker GPU运行时配置
- 确保已安装NVIDIA Docker运行时,执行
docker run --rm --gpus all nvidia/cuda:11.8.0-base-ubuntu22.04 nvidia-smi测试基础GPU容器是否能识别GPU。 - 在Docker Compose配置中,若使用Docker 20.10+,
deploy.resources.reservations.devices配置有效;若版本较低,需添加runtime: nvidia字段:i2v-neural: image: semitechnologies/img2vec-pytorch:resnet50 environment: ENABLE_CUDA: '1' CUDA_VISIBLE_DEVICES: '0' # 指定具体GPU编号,避免"all"可能的兼容性问题 runtime: nvidia deploy: resources: reservations: devices: - driver: nvidia count: 1 capabilities: [gpu]
- 确保已安装NVIDIA Docker运行时,执行
检查环境变量与容器内CUDA兼容性
- 确认容器内PyTorch版本支持主机的CUDA驱动版本,可手动进入容器执行
python -c "import torch; print(torch.version.cuda); print(torch.cuda.is_available())"验证。 - 调整环境变量
CUDA_CORE(若镜像支持),明确指定cuda:0而非依赖自动检测。
- 确认容器内PyTorch版本支持主机的CUDA驱动版本,可手动进入容器执行
确认主机驱动与容器CUDA版本匹配
RTX3060需要CUDA 11.0及以上版本的驱动,检查主机驱动版本(nvidia-smi查看),确保容器内的CUDA版本不超过驱动支持的最高版本。测试容器内GPU可用性
手动启动容器并执行命令,排除配置问题:docker run --rm --gpus all semitechnologies/img2vec-pytorch:resnet50 python -c "import torch; print(torch.cuda.is_available())"若输出
True,说明容器本身可访问GPU,问题出在Docker Compose配置或镜像启动逻辑。
内容的提问来源于stack exchange,提问作者user2741831
相关产品推荐
相关产品推荐

