You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AzureML批量端点调用失败:azureml-dataprep版本兼容问题求助

问题描述

背景

将图片上传至Blob Storage的文件夹,创建指向该文件夹的AzureML数据资产(Data asset),并将其作为输入调用AzureML批量推理端点。核心代码如下:

my_data = Data(
    path=DATA_ASSET_PATH,
    type=AssetTypes.URI_FOLDER,
    name=f"notebook_{temp_uuid}",
)
ml_client.data.create_or_update(my_data)

job = ml_client.batch_endpoints.invoke(
    endpoint_name=cfg["batch_endpoint_name"],
    inputs={
        "input": Input(path=DATA_ASSET_PATH, type=AssetTypes.URI_FOLDER),
    },
    output_path={
        "score": Input(path=os.path.join(DATA_ASSET_PATH, "output"), type=AssetTypes.URI_FOLDER)
    },
    output_file_name="predictions.csv",
    params_override=[
        {"mini_batch_size": "20"},
        {"compute.instance_count": "1"},
    ],
)

端点配置使用项目的Docker镜像环境,该镜像通过Poetry管理依赖。

异常现象

端点调用的执行环境会影响数据处理,出现以下错误:

UserErrorException:     
Message: Failed to load dataset definition with azureml-dataprep==4.12.9. 
Please install the latest version with "pip install -U azureml-dataprep".   
InnerException None     ErrorResponse { "error": { "code": "UserError", "message": "Failed to load dataset definition with azureml-dataprep==4.12.9. Please install the latest version with \"pip install -U azureml-dataprep\"." } }

不同场景测试结果:

  • 本地笔记本执行上述代码:运行成功
  • FastAPI(部署于App Service)或Azure Function(本地/云端运行)执行:触发上述错误
  • 仅创建Data asset,手动在AzureML Studio触发任务:运行成功

将Docker环境中的azureml-dataprep升级至5.1.4(已在poetry.lock中配置)后,错误提示变为azureml-dataprep==4.12.10,且堆栈跟踪显示使用Python 3.8,但Docker镜像实际使用Python 3.10。


排查与修复方案

1. 对齐执行环境的依赖版本

本地笔记本、App Service/Azure Function、AzureML批量推理环境三者的azureml-dataprep及AzureML SDK版本可能存在差异:

  • 检查App Service/Azure Function的依赖清单(pyproject.toml或生成的requirements.txt),确认是否安装了与Docker镜像一致的azureml-dataprep和azure-ai-ml版本。
  • 本地运行成功是因为其依赖版本与AzureML批量环境兼容,而App Service/Azure Function的旧版本依赖会导致创建的Data asset携带不兼容的数据集定义,触发服务端版本校验错误。

2. 强制指定依赖版本

在App Service/Azure Function的Poetry配置中,明确锁定与Docker镜像匹配的依赖版本:

# pyproject.toml
[tool.poetry.dependencies]
python = "3.10.*"
azure-ai-ml = "<与AzureML环境匹配的版本>"
azureml-dataprep = "5.1.4"
# 其他必要依赖...

重新生成poetry.lock并部署,确保客户端环境依赖与服务端完全对齐。

3. 验证Docker镜像构建正确性

确认Docker镜像在构建时确实安装了指定版本的依赖:

  • 在Dockerfile中添加版本校验命令,构建时查看输出确认版本:
RUN poetry run pip show azureml-dataprep && poetry run python --version
  • 检查Dockerfile是否正确执行poetry install --no-root,未出现覆盖依赖版本的操作。

4. 调整Data asset创建流程

错误中的“dataset definition”由客户端创建Data asset时生成,可能携带客户端环境的版本信息。可调整流程:

  • 提前在AzureML环境(Studio/CLI/笔记本)创建好Data asset,App Service/Azure Function调用端点时直接使用该资产名称,而非Blob路径:
job = ml_client.batch_endpoints.invoke(
    endpoint_name=cfg["batch_endpoint_name"],
    inputs={
        "input": Input(type=AssetTypes.URI_FOLDER, path="<已创建的Data asset名称>"),
    },
    # 其他参数...
)

让批量推理环境使用自身依赖处理数据集,避免客户端与服务端版本冲突。

5. 检查批量端点的计算环境配置

堆栈跟踪显示使用Python 3.8,可能是端点绑定的计算集群未使用自定义Docker镜像:

  • 登录AzureML Studio,查看批量端点的部署配置,确认使用的是自定义Docker镜像,而非AzureML默认的Curated环境。
  • 若使用了Curated环境,切换为自定义镜像,确保计算实例运行的Python版本和依赖与镜像一致。

内容的提问来源于stack exchange,提问作者jerorx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 11:42:19