You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure ML自定义CUDA环境下PyTorch CUDA设备不可用问题求助

自定义Azure ML环境中CUDA任务运行失败问题

过去一周,我尝试在Azure ML Studio中创建Python实验,任务为使用搭载CUDA 11.6的自定义环境训练PyTorch(1.12.1)神经网络以实现GPU加速。但执行张量移动操作时触发Runtime Error:

device = torch.device("cuda")
test_tensor = torch.rand((3, 4), device = "cpu")
test_tensor.to(device)

错误信息:

CUDA error: all CUDA-capable devices are busy or unavailable
CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.
For debugging consider passing CUDA_LAUNCH_BLOCKING=1.

已尝试的排查操作

  • 设置CUDA_LAUNCH_BLOCKING=1,结果未改变
  • 检查CUDA可用性,执行以下代码:
    print(f"Is cuda available? {torch.cuda.is_available()}")
    print(f"Which is the current device? {torch.cuda.current_device()}")
    print(f"How many devices do we have? {torch.cuda.device_count()}")
    print(f"How is the current device named? {torch.cuda.get_device_name(torch.cuda.current_device())}")
    
    输出结果完全正常:
    Is cuda available? True
    Which is the current device? 0
    How many devices do we have? 1
    How is the current device named? Tesla K80
    
  • 尝试降级或更换CUDA、Torch及Python版本,未解决错误
  • 排查发现错误仅在使用自定义环境时出现,使用Azure托管环境时脚本可正常运行,但因脚本依赖OpenCV等库,必须使用自定义Dockerfile创建环境

自定义Dockerfile内容

FROM mcr.microsoft.com/azureml/aifx/stable-ubuntu2004-cu116-py39-torch1121:biweekly.202301.1


USER root
RUN apt update
# Necessary dependencies for OpenCV
RUN apt install ffmpeg libsm6 libxext6 libgl1-mesa-glx -y 

RUN pip install numpy matplotlib pandas opencv-python Pillow scipy tqdm mlflow joblib onnx ultralytics
RUN pip install 'ipykernel~=6.0' \
                'azureml-core' \
        'azureml-dataset-runtime' \
                'azureml-defaults' \
        'azure-ml' \
        'azure-ml-component' \
                'azureml-mlflow' \
                'azureml-telemetry' \
        'azureml-contrib-services'

COPY --from=mcr.microsoft.com/azureml/o16n-base/python-assets:20220607.v1 /artifacts /var/
RUN /var/requirements/install_system_requirements.sh && \
    cp /var/configuration/rsyslog.conf /etc/rsyslog.conf && \
    cp /var/configuration/nginx.conf /etc/nginx/sites-available/app && \
    ln -sf /etc/nginx/sites-available/app /etc/nginx/sites-enabled/app && \
    rm -f /etc/nginx/sites-enabled/default
ENV SVDIR=/var/runit
ENV WORKER_TIMEOUT=400
EXPOSE 5001 8883 8888

注:上述COPY语句中的代码复制自Azure预定义的托管环境,即使直接使用托管环境的Dockerfile未作任何修改,仍会出现相同错误。

问题

如何在自定义环境中运行CUDA任务?这是否可行?

我曾尝试寻找解决方案,但未找到遇到相同问题的用户,也未在微软文档中找到相关咨询渠道,希望能得到帮助。


内容的提问来源于stack exchange,提问作者carlosmhd27

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 01:50:29