You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Azure SDKv2的Pipeline作业中访问私有Blob容器?

解决Azure ML Pipeline访问私有Blob容器的配置方法

场景:在Azure ML Pipeline作业中使用自定义Input访问私有Blob容器,已为作业运行的计算集群托管标识配置了存储账户(含目标私有容器)及机器学习工作区的Owner和Contributor权限,但作业无法访问该容器内的数据。当前自定义Input配置如下:

Input(
    type="uri_folder",
    path="wasbs://container_name@storage_account.blob.core.windows.net/",
    mode="ro_mount"
),

核心配置步骤

  • 启用作业的托管身份认证:因已为计算集群托管标识配置存储权限,需在Pipeline作业(或组件)中明确指定使用该身份,替代默认的None配置。
  • 补充数据访问权限:仅配置工作区Owner/Contributor权限不足,需确保计算集群托管标识拥有存储容器的Storage Blob Data Reader(或更高)数据访问权限。

修改后的代码示例

1. 配置作业身份

在创建Command组件(或Pipeline作业)时,指定使用计算集群的托管身份:

from azure.ai.ml import command, Input, ManagedIdentityConfiguration

# 保留你的自定义Input配置
custom_input = Input(
    type="uri_folder",
    path="wasbs://container_name@storage_account.blob.core.windows.net/",
    mode="ro_mount"
)

# 创建作业时启用托管身份认证
job = command(
    command="ls ${{inputs.custom_data}}",  # 示例命令验证数据访问
    inputs={"custom_data": custom_input},
    environment="azureml://registries/azureml/environments/sklearn-1.1/versions/4",
    compute="cpu-cluster",
    # 关键配置:使用计算集群的托管身份
    identity=ManagedIdentityConfiguration()
)

2. 权限验证与补充

进入存储账户的IAM页面,添加角色分配:

  • 选择角色:Storage Blob Data Reader
  • 分配对象:计算集群的系统分配托管标识(名称与集群名称一致)

额外优化建议

若需简化后续配置,可将私有Blob容器关联为Azure ML Datastore:

  1. 创建Datastore时指定使用计算集群托管身份认证
  2. Input路径直接引用Datastore:azureml://datastores/your_datastore_name/paths/
    此时作业会自动继承Datastore的认证方式,无需重复配置身份。

内容的提问来源于stack exchange,提问作者Timbus Calin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 20:02:48