SageMaker GPU Docker实例本该含CUDA却缺失问题求助
问题:SageMaker Processing无法启用GPU运行PyTorch任务
背景与代码配置
我尝试使用SageMaker Processing在GPU上运行PyTorch模型的训练与评估,代码配置如下:
from sagemaker.processing import ScriptProcessor role = 'arn:aws:iam.../service-role/SageMaker-MySageMakerComputeRole' script_processor = ScriptProcessor( command=["python3"], image_uri='...dkr.ecr.us-west-1.amazonaws.com/sagemaker-processing-container:latest', role=role, instance_count=1, instance_type="ml.m5.xlarge", )
已尝试的镜像配置
参考AWS官方镜像列表,我基于两种PyTorch GPU镜像构建自定义容器:
# 第一种尝试的镜像 FROM 763104351884.dkr.ecr.us-west-1.amazonaws.com/pytorch-training:1.13.1-gpu-py39-cu117-ubuntu20.04-sagemaker # 第二种尝试的镜像 FROM 763104351884.dkr.ecr.us-west-1.amazonaws.com/pytorch-training:1.13.1-gpu-py39-cu117-ubuntu20.04-ec2 WORKDIR / RUN apt-get update && apt-get install -y git ...
报错详情
任务日志持续返回如下错误:
2023-06-19T01:48:46.062-07:00 RuntimeError: ('Found no NVIDIA driver on your system. Please check that you 2023-06-19T01:48:46.062-07:00 have an NVIDIA GPU and installed a driver from 2023-06-19T01:48:46.062-07:00 http://www.nvidia.com/Download/index.aspx', "This image doesn't seem to support 2023-06-19T01:48:46.062-07:00 current EC2 instance type, please check release notes for supported EC2 instance 2023-06-19T01:48:46.062-07:00 type")
同时在Processing脚本中执行torch.cuda.is_available()返回False,确认当前环境无CUDA支持。
疑问
我理解官方GPU镜像应该预装CUDA相关依赖,请问是否存在镜像选择错误,或是遗漏了关键配置步骤?如何解决才能让SageMaker Processing启用GPU工作负载?
内容的提问来源于stack exchange,提问作者Sticky
相关产品推荐
相关产品推荐

