You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

自定义Azure环境训练时出现TensorFlow的ModuleNotFoundError

问题:自定义AzureML环境找不到TensorFlow依赖

我有一个TensorFlow脚本,使用ACR镜像mcr.microsoft.com/azureml/curated/tensorflow-2.7-ubuntu20.04-py38-cuda11-gpu:28(AzureML-tensorflow-2.7-ubuntu20.04-py38-cuda11-gpu)训练时完全正常。为添加额外依赖,我基于该镜像创建了自定义环境,镜像构建成功,但训练任务执行时提示找不到tensorflow。

注:在conda.yml中我注释掉了原镜像已包含的包,不注释的话镜像创建会失败。

conda.yml内容

name: keras-env
channels:
  - conda-forge
dependencies:
  - python=3.8
  - pip=20.2.4
  - pip:
    #- protobuf~=3.20
    #- numpy~=1.21.0
    #- tensorflow-gpu~=2.7.0
    #- matplotlib~=3.5.0
    #- azureml-mlflow==1.51.0
    #- horovod[tensorflow-gpu]~=0.23.0
    - azureml.core
    - keras
    - mlflow
    - pyarrow
    - idx2numpy
    - scikit-learn

环境构建代码

import os
from azure.ai.ml.entities import Environment

custom_env_name = "stb-dist-keras-env"
dependencies_dir = "./"

from azure.ai.ml import MLClient
from azure.identity import DefaultAzureCredential

ml_client = MLClient(
    DefaultAzureCredential(), 'a', 'ab', 'abc'
)

job_env = Environment(
    name=custom_env_name,
    description="Custom environment distributed environment",
    conda_file=os.path.join(dependencies_dir, "conda.yml"),
    image="mcr.microsoft.com/azureml/curated/tensorflow-2.7-ubuntu20.04-py38-cuda11-gpu:28"
)
job_env = ml_client.environments.create_or_update(job_env)

print(
    f"Environment with name {job_env.name} is registered to workspace, the environment version is {job_env.version}"
)

错误日志

Traceback (most recent call last): File "train.py", line 18, in
import tensorflow as tf ModuleNotFoundError: No module named 'tensorflow'

补充信息

  • 训练脚本train.py使用的是Horovod官方的TensorFlow2-Keras-MNIST示例代码,两次训练用的是同一个脚本。
  • 自定义环境构建成功截图:自定义环境构建成功
  • 原MCR镜像训练成功截图:原镜像训练成功
  • 自定义环境训练失败截图:自定义环境训练失败

任务启动代码

from azure.ai.ml import command, MpiDistribution

job = command(
    code="./",  # local path where the code is stored
    command="python train.py --epochs ${{inputs.epochs}}",
    inputs={"epochs": 1},
    #environment="AzureML-tensorflow-2.7-ubuntu20.04-py38-cuda11-gpu@latest",
    environment="stb-dist-keras-env@latest",
    compute="",
    instance_count=1,
    distribution=MpiDistribution(process_count_per_instance=1),
    display_name="tensorflow-mnist-distributed-horovod-example"
    # experiment_name: tensorflow-mnist-distributed-horovod-example
    # description: Train a basic neural network with TensorFlow on the MNIST dataset, distributed via Horovod.
)

内容的提问来源于stack exchange,提问作者webber

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 14:12:40