Azure ML自定义CUDA环境下PyTorch CUDA设备不可用问题求助
自定义Azure ML环境中CUDA任务运行失败问题
过去一周,我尝试在Azure ML Studio中创建Python实验,任务为使用搭载CUDA 11.6的自定义环境训练PyTorch(1.12.1)神经网络以实现GPU加速。但执行张量移动操作时触发Runtime Error:
device = torch.device("cuda") test_tensor = torch.rand((3, 4), device = "cpu") test_tensor.to(device)
错误信息:
CUDA error: all CUDA-capable devices are busy or unavailable CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect. For debugging consider passing CUDA_LAUNCH_BLOCKING=1.
已尝试的排查操作
- 设置
CUDA_LAUNCH_BLOCKING=1,结果未改变 - 检查CUDA可用性,执行以下代码:
输出结果完全正常:print(f"Is cuda available? {torch.cuda.is_available()}") print(f"Which is the current device? {torch.cuda.current_device()}") print(f"How many devices do we have? {torch.cuda.device_count()}") print(f"How is the current device named? {torch.cuda.get_device_name(torch.cuda.current_device())}")Is cuda available? True Which is the current device? 0 How many devices do we have? 1 How is the current device named? Tesla K80 - 尝试降级或更换CUDA、Torch及Python版本,未解决错误
- 排查发现错误仅在使用自定义环境时出现,使用Azure托管环境时脚本可正常运行,但因脚本依赖OpenCV等库,必须使用自定义Dockerfile创建环境
自定义Dockerfile内容
FROM mcr.microsoft.com/azureml/aifx/stable-ubuntu2004-cu116-py39-torch1121:biweekly.202301.1 USER root RUN apt update # Necessary dependencies for OpenCV RUN apt install ffmpeg libsm6 libxext6 libgl1-mesa-glx -y RUN pip install numpy matplotlib pandas opencv-python Pillow scipy tqdm mlflow joblib onnx ultralytics RUN pip install 'ipykernel~=6.0' \ 'azureml-core' \ 'azureml-dataset-runtime' \ 'azureml-defaults' \ 'azure-ml' \ 'azure-ml-component' \ 'azureml-mlflow' \ 'azureml-telemetry' \ 'azureml-contrib-services' COPY --from=mcr.microsoft.com/azureml/o16n-base/python-assets:20220607.v1 /artifacts /var/ RUN /var/requirements/install_system_requirements.sh && \ cp /var/configuration/rsyslog.conf /etc/rsyslog.conf && \ cp /var/configuration/nginx.conf /etc/nginx/sites-available/app && \ ln -sf /etc/nginx/sites-available/app /etc/nginx/sites-enabled/app && \ rm -f /etc/nginx/sites-enabled/default ENV SVDIR=/var/runit ENV WORKER_TIMEOUT=400 EXPOSE 5001 8883 8888
注:上述COPY语句中的代码复制自Azure预定义的托管环境,即使直接使用托管环境的Dockerfile未作任何修改,仍会出现相同错误。
问题
如何在自定义环境中运行CUDA任务?这是否可行?
我曾尝试寻找解决方案,但未找到遇到相同问题的用户,也未在微软文档中找到相关咨询渠道,希望能得到帮助。
内容的提问来源于stack exchange,提问作者carlosmhd27
相关产品推荐
相关产品推荐

