You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure ML计算集群User Assigned Identity配置问题咨询

解决Azure ML Compute Cluster继承身份或使用当前用户身份的问题

一、让Compute Cluster使用指定的用户分配身份(与Compute Instance同身份)

Compute Cluster不会自动继承Compute Instance的身份,必须显式在创建时指定完整的用户分配身份ARM ID,而非仅实例化空的UserIdentityConfiguration()。

1. 获取用户分配身份的ARM ID

在你的Compute Instance中,通过以下命令获取目标身份的完整ID:

az identity list --query "[?name=='<你的用户分配身份名称>'].id" -o tsv

也可通过Azure SDK从当前Compute Instance的元数据中获取:

import requests

metadata_url = "http://169.254.169.254/metadata/identity/info?api-version=2018-02-01"
headers = {"Metadata": "true"}
response = requests.get(metadata_url, headers=headers).json()
identity_id = response['identityIds'][0]  # 若实例绑定多个身份,需选择对应项

2. 使用Azure SDK v2正确配置集群身份

以Python SDK为例,创建集群时需将身份ID传入UserIdentityConfiguration的user_assigned_identities参数:

from azure.mgmt.machinelearningservices import MachineLearningServicesClient
from azure.mgmt.machinelearningservices.models import AmlCompute, UserIdentityConfiguration, ComputeResource

# 初始化ML客户端(已通过Compute Instance身份认证)
ml_client = MachineLearningServicesClient.from_config()

# 替换为你的资源信息和身份ID
resource_group = "<资源组名称>"
workspace_name = "<工作区名称>"
cluster_name = "<集群名称>"
identity_id = "<获取到的用户分配身份ARM ID>"

# 构建集群配置
compute_config = AmlCompute(
    properties={
        "vmSize": "Standard_DS3_v2",
        "scaleSettings": {"maxNodeCount": 4},
        "identity": UserIdentityConfiguration(
            user_assigned_identities={identity_id: {}}
        )
    }
)

# 创建/更新集群
ml_client.compute.begin_create_or_update(
    resource_group_name=resource_group,
    workspace_name=workspace_name,
    compute_name=cluster_name,
    parameters=ComputeResource(properties=compute_config)
).result()

这样创建的集群,其yml配置中的identity字段会正确填充用户分配身份信息。

二、让Compute Cluster作业使用当前az login用户身份

如果希望集群运行作业时使用执行代码的az login用户身份,需启用身份传递(Identity Passthrough),无需配置集群自身的用户分配身份,而是在提交作业时指定用户身份:

1. 提交作业时配置用户身份

from azure.ai.ml import command, MLClient
from azure.ai.ml.entities import UserIdentity

# 初始化ML客户端(使用az login的用户身份认证)
ml_client = MLClient.from_config()

# 定义作业
job = command(
    code="./src",
    command="python train.py",
    environment="AzureML-sklearn-0.24-ubuntu18.04-py37-cpu@latest",
    compute="<集群名称>",
    identity=UserIdentity()  # 指定使用提交作业的用户身份
)

# 提交作业
returned_job = ml_client.jobs.create_or_update(job)

作业运行时会自动使用提交者的身份访问Azure服务,前提是该用户拥有目标服务(CosmosDB、Databricks)的RBAC权限。

三、排查集群yml中identity为null的问题

  • 检查SDK代码:确保没有仅实例化空的UserIdentityConfiguration(),必须传入user_assigned_identities字典,键为身份ARM ID,值为空字典。
  • 验证身份权限:确保Compute Instance使用的用户分配身份拥有以下权限:
    • Microsoft.MachineLearningServices/workspaces/computes/write(创建/更新集群)
    • Microsoft.ManagedIdentity/userAssignedIdentities/assign/action(将身份分配给集群)
  • 确认身份ID正确性:必须使用完整的ARM ID,格式为/subscriptions/<订阅ID>/resourceGroups/<资源组>/providers/Microsoft.ManagedIdentity/userAssignedIdentities/<身份名称>

内容的提问来源于stack exchange,提问作者BeGreen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 03:05:18